Synthetic Data
Feature Guide — Building Realistic, Vendor-Shaped Test Data
Module: Discovery Module › Synthetic Data | Last updated: August 2026
Contents
- Overview
- Finding Your Way Around
- Projects — New, Open and Resume
- Step 1 — Setup
- Step 2 — Source Model
- Step 3 — Generation Strategy
- Step 4 — Intelligent Mapping
- Step 5 — Review Mappings
- Step 6 — Preview
- Step 7 — Execution
- Quick Reference
1. Overview
Synthetic Data builds realistic, made-up datasets shaped like a real source system — the right tables, the right columns, sensible-looking values, and working relationships between them — without using any real, sensitive data. It's ideal for testing, demos, training, and development where you need data that looks real but carries no privacy or security risk.
It works from the same vendor metadata catalog behind the rest of the Discovery Module (see Catalog Admin). You pick the objects you want, choose how each column's values should be produced, preview the result, and then generate a full dataset — optionally writing it straight into a connected database.
The whole process is a guided, seven-step wizard, so you're walked through it from start to finish.
What you can do here
- Choose a set of related tables from a vendor catalog as your scope.
- Control how each column's values are generated, or let the system decide.
- Preview a small sample before committing.
- Validate the plan, then generate a full dataset.
- Write the results into a target database, or just generate and count.
- Save projects and pick them up again later.
WHERE TO FIND IT — Synthetic Data is a single workbench page with a seven-step wizard across the top. It's reached from its own left-nav entry.
YOU NEED THE "USE SYNTHETIC DATA" PERMISSION — The page and its menu entry only appear for users who've been granted access to the Synthetic Data workbench.
2. Finding Your Way Around
The workbench is a wizard with seven steps, shown as a strip across the top:
Setup → Source Model → Generation Strategy → Intelligent Mapping → Review Mappings → Preview → Execution
Move through them with the Back and Next buttons at the bottom, which stay visible as you scroll. You can also click a step in the strip to jump straight to it. The Next button on Setup only enables once you've given the project a name and added at least one object, so you can't skip the essentials.
Each step is covered in its own section below. The final step, Execution, runs four phases in sequence — Generation plan, Validation, Generate and Results — all on the one screen.
3. Projects — New, Open and Resume
Your work is saved as a project, so you can build a dataset over time and return to it. Two buttons sit above the wizard:
- New project clears everything and starts fresh on Setup.
- Open project lists your saved projects — each showing its vendor, who created it, when it was last updated, and its status — so you can reopen one.
When you reopen a project, the workbench drops you back where you left off: a generated project reopens on Execution, a mapped project on Review Mappings, and anything earlier on Setup.
PROJECTS SAVE AUTOMATICALLY PART-WAY THROUGH — A project is created and saved once you reach the Intelligent Mapping step, because the later steps work against a saved project. Up to that point your work lives in the current session.
4. Step 1 — Setup
Setup is where you name the project and build its scope — the set of objects you want data for.
At the top, fill in the project Name (required), an optional Description, and the Domain Layer the project belongs to.
Below that, the screen is split in two. On the left is the catalog search; on the right is the scope canvas.
Finding objects (left) — Narrow the catalog with the filters, then search:
| Filter | What it does |
|---|---|
| Vendor System | The source system to pull objects from (for example, SAP or Baan LN). |
| Version | The specific version of that vendor (objects are version-specific). |
| Business object | Limit to objects in one business-object group. |
| Industry | Limit to a specific industry solution, or core/cross-app objects. |
| Search objects | Free-text search across code, name, description, module and package. |
Results show each object's code, business name, description, module/package, classification and relationship count. Click Add on a row to place it on the canvas.
Building scope (right) — The canvas shows the objects you've added and the relationships between them. The first object you add becomes the root (marked accordingly). You can:
- Right-click a node → Expand related to automatically pull in the tables it's linked to (its foreign-key neighbours), going deeper if you choose.
- ✕ to remove an object.
- Drag nodes to arrange them.
A counter shows how many objects are in scope. If an expansion brings in a very large number of objects, the canvas caps the amount and invites you to go deeper deliberately.
CHANGING VENDOR OR VERSION CLEARS THE SCOPE — Because objects belong to a specific vendor version, switching either one resets the canvas. Settle on the vendor and version before you build scope.
5. Step 2 — Source Model
Source Model is a read-only review of the objects in your scope. The left rail lists every object; select one to see its columns on the right — column name, data type, length, whether it's nullable, its key role (PK/FK), and its description.
Use this step to confirm you've got the right tables and to understand their shape before deciding how to generate their values. There's nothing to change here — when you're happy, continue to the next step.
6. Step 3 — Generation Strategy
This step decides how values are produced. It has two parts.
Strategy cards (top) — Pick the overall approach. Each card describes a strategy and the data domains it covers; the first is selected by default (the native source-model strategy, which generates values shaped by the catalog's own types and names).
Per-field generators (below) — By default every column uses the Generic (auto) generator, which produces sensible values from each column's data type and name (emails, dates, amounts, codes, and so on). You only need to step in where you want tighter control. Choose an object from the picker, then click Edit on any column to assign a specific generator:
| Generator | What it produces |
|---|---|
| Generic (auto) | Automatic, type- and name-aware values. The default; no setup needed. |
| Choice | A value picked from a list you supply. |
| Sequence | Sequential values (for example, running numbers). |
| Pool Sample | Values sampled from a defined pool. |
| Date Range | Dates within a range you set. |
| Expression | A value from a small formula you write, for full control. |
The Object dropdown — Your scope can hold several objects, and generators are assigned one object at a time. The Object picker chooses which object's columns the table below shows, so switch it to set up each object in turn. Your assignments are remembered per object, so moving between them never loses your choices.
Each generator has its own simple parameters in the edit dialog. If you're assigning the same column across several objects in scope, you can tick an option to apply it to that column in every scope object at once. To undo an override, open the column and reset it to the default.
EXPRESSIONS ARE CHECKED AS YOU TYPE — The Expression generator lets you write a short formula. It's validated live and certain unsafe keywords are blocked, so you'll know immediately whether an expression is valid before you save it.
7. Step 4 — Intelligent Mapping
Intelligent Mapping automatically works out how your objects' columns line up, matching each source object to its target and scoring how confident it is. Click Run intelligent mapping to analyse the scope.
You can adjust two thresholds before running — the minimum table score and minimum column score — to control how strict the matching is.
What the thresholds mean — Every candidate match gets a score from 0 to 100% for how confident the system is. The two thresholds set the cut-off for keeping a match:
- Min table score — how confident the match between two tables must be before the pair is kept. Raise it to see only strong table matches; lower it to surface weaker ones.
- Min column score — how confident the match between two columns must be before those columns count as mapped. Raise it for stricter column matching; lower it to accept looser pairings.
Higher thresholds give fewer but more certain matches; lower thresholds give more matches but more to review.
After a run, summary cards show the number of matches and the average confidence, and a card per table pair shows its confidence, mapped column count, and how many source and target columns went unmatched.
If nothing scores above the table threshold, lower it and run again.
AI PROPOSES, YOU DECIDE — Mapping only suggests how things line up. Nothing is committed here — you review and approve the suggestions in the next step.
8. Step 5 — Review Mappings
Here you review what Intelligent Mapping proposed and decide what to keep. Each table pair is shown with three columns: its source columns, the matched target columns, and any that need attention (unmatched source columns, flagged so you can see the gaps).
| Column | What it lists |
|---|---|
| Source columns | Every column in the source table — the full set the mapping considered. |
| Matched target columns | The target columns that were successfully paired to a source column. |
| Needs attention | Source columns that found no confident match, each tagged unmatched — the gaps to check or map by hand. |
Use the Accepted / Rejected toggle on each pair to include or exclude it, then click Apply accepted mappings to save your decisions. A confirmation tells you how many were applied.
9. Step 6 — Preview
Preview generates a small 100-row sample of a chosen object so you can eyeball the shape of the output before committing to a full run. Pick the object, click Generate preview, and the rows appear in a table.
Each column header shows which generator produced it, so you can confirm your Generation Strategy choices are doing what you expect. This is purely a look — nothing is saved or written — so preview freely and go back to adjust generators if something looks off.
10. Step 7 — Execution
Execution is the final step, and it runs in four phases on one screen.
1. Generation plan — Shows the order tables will be generated in (arranged so parent tables come before the children that depend on them), and, per column, which generator will run and what it depends on. It's a preview of exactly what's about to happen.
2. Validation — Click Validate project to check the plan for problems. Findings are grouped by severity (Critical, High, Medium, Low, Info). Any critical findings block generation until they're resolved, so this is your safety check before producing data.
3. Generate — Set the options, then click Generate dataset:
| Option | What it does |
|---|---|
| Rows per table | How many rows to produce for each table. |
| Write to target | On: write the rows into a database. Off: just generate and count them. |
| Target source / Target schema | The connection and schema to write into (required when writing to a target). |
| Table prefix | An optional prefix for the created table names. |
| Provenance columns | Adds columns recording how each row was produced. |
| Referential integrity | Keeps foreign keys valid — children reference real parent keys. |
| Avg. children / parent | With referential integrity on, roughly how many child rows per parent. |
| Strict length | On: fail a table if a value exceeds a column's declared length. Off: truncate to fit and report. |
YOU CAN GENERATE WITHOUT WRITING — Turn Write to target off to generate and count rows without touching any database — useful for a dry run. When it's on, you must choose a target source and schema, or Generate stays disabled. Tables are written drop-and-recreate, so an existing table of the same name is replaced.
4. Results — After a run, a summary shows the run number and status, total rows produced, and a card per table with how many rows were generated and written (and a reason if any table failed). If you reopen a project that's already been generated, its last results are restored here.
11. Quick Reference
A fast lookup for the most common actions.
| I want to… | Do this |
|---|---|
| Start a new dataset | New project → Setup |
| Reopen earlier work | Open project → pick from the list |
| Choose which tables to generate | Setup → filter/search the catalog → Add to the canvas |
| Pull in related tables automatically | Setup → right-click a node → Expand related |
| Check a table's columns | Source Model → select the object |
| Control how a column is generated | Generation Strategy → Edit the column → pick a generator |
| Auto-match columns | Intelligent Mapping → Run intelligent mapping |
| Keep or drop proposed mappings | Review Mappings → Accepted/Rejected → Apply accepted mappings |
| See a small sample first | Preview → Generate preview |
| Check for problems before generating | Execution → Validate project |
| Produce the full dataset | Execution → set options → Generate dataset |
| Generate without writing to a database | Execution → turn off Write to target → Generate dataset |
Source: OnCoor Discovery Module product documentation, written for end users. For the latest screens and options, always refer to the in-app interface.