Skip to main content

Discovery Workbench

Feature Guide — Profiling a Source, Finding Issues, and Promoting Fixes​

Module: Discovery Module › Discovery Workbench  |  Last updated: August 2026


Contents​

  1. Overview
  2. Finding Your Way Around
  3. Starting a Session
  4. Loading a Sample
  5. Deterministic vs AI Mode
  6. Profile Criteria and Checks
  7. Running the Profile
  8. Findings and Promoting Fixes
  9. Duplicate Detection
  10. External Validation
  11. Modeling in the Platform
  12. Reports
  13. Reference Data and Providers
  14. Quick Reference

1. Overview​

The Discovery Workbench lets you profile a source table — take a real sample of its data, check it against a set of quality and governance rules, score how ready it is, spot duplicate records, and then promote the issues it finds into draft fixes elsewhere in OnCoor. It's how an analyst gets from "here's a table I don't fully trust" to "here are the specific problems, and here are drafted fixes to review."

The workbench is built around a session — one session profiles one table. You start a session, load a sample, run a profile, review the findings and any duplicate clusters, and promote what matters. Sessions are saved, so you can resume them later.

A guiding principle runs through the whole tool: AI proposes, a human disposes. The workbench detects issues with built-in, deterministic rules; AI (when you turn it on) only explains findings, matches similar records, and drafts fixes. Nothing is changed in your data, and every promotion lands as a draft for you to review — never an automatic change.

What you can do here

  • Profile a source table against built-in and custom quality checks.
  • See a readiness score and plain-language findings.
  • Find duplicate records and review the clusters.
  • Validate values against reference lists and external services (email, phone, postal, and more).
  • Promote findings to a draft Rule Pilot rule or Cleansing transformation, and duplicates to a De-Duplication rule.
  • Produce a report as PDF, Word or Excel.

WHERE TO FIND IT — The Discovery Workbench has its own left-nav entry. It only appears for users granted the Discovery Explorer role (the Use Discovery Workbench permission).

NOTHING HERE CHANGES YOUR DATA — Profiling only reads a sample. Detection is deterministic and every promotion creates a draft artifact for review. You're always in control of what actually gets applied.


2. Finding Your Way Around​

When you open the workbench you're on the start screen, where you begin a new session or resume a saved one. There are also two side doors here — Reports and Reference & Providers — that don't need an active session (covered near the end of this guide).

Once a session is active, the screen has three parts:

  • A toolbar across the top with the table name, the session status, and the main actions (Load sample, mode toggle, Criteria, Run Profile, DeDup).
  • A sample grid on the left showing the sampled rows.
  • A profile rail on the right showing the readiness score, the narrative, and tabs for Findings and Duplicate clusters.

Use ← Return to Workbench (top right) to go back to the start screen, and Report to view the current session's report.


3. Starting a Session​

A session profiles one table, and there are two ways to choose it — pick the mode with the toggle at the top:

ModeWhat it's for
From catalogBrowse the vendor metadata catalog (the same one behind Catalog Admin). You get rich detail — code, business name, description, module/package, class and relationship count — because the table is described in the catalog.
Direct from sourceBrowse a connection's live physical tables directly, with no catalog needed. You choose a connection and pick a table from its schema.

In catalog mode, choose the source connection, then narrow the catalog with the Vendor System, Version and Business object filters and the Search objects box. Click a row to select the object.

In direct mode, choose the connection and pick a table from the list.

Either way, click Start session once you've selected something. Both modes need a source connection.

Resuming — If you have saved sessions, a Resume a session list appears on the start screen, showing each session's table, status, mode, when it was created, and its readiness score. Click one to reopen it with its saved findings, clusters and score restored. (The row sample isn't stored, so you'll re-load it if you want to see rows again.)


4. Loading a Sample​

With a session open, click Load sample to pull a sample of rows (up to 1,000) from the table into the grid. The sample is what the profile and duplicate detection run against. Click Reload sample any time to refresh it.

Loading a sample also reveals the DeDup on panel (see Duplicate Detection) and populates the column pickers used by the criteria editor.

SYSTEM COLUMNS ARE SET ASIDE — Behind-the-scenes provenance columns (for example, loader run IDs and timestamps) aren't part of the real data model, so they're hidden from profiling and duplicate detection.


5. Deterministic vs AI Mode​

A toggle in the toolbar switches between two modes, and it governs all three AI-capable actions — Run Profile, DeDup and Promote:

ModeWhat happens
DeterministicEverything runs with built-in rules only. Fast, fully offline, and no AI is called. This is the default.
✦ AIAdds AI assistance: Run Profile writes plain-English explanations of each finding; DeDup also matches records that mean the same thing (not just textual duplicates); Promote auto-drafts the rule SQL, cleansing expression or dedup keys for you to review.

Buttons marked with a ✦ are the ones that call AI. You can switch modes at any time between runs.

DETECTION IS ALWAYS DETERMINISTIC — Even in AI mode, the detection of issues and duplicates uses built-in rules. AI only explains, matches similar records, and drafts fixes — it never decides what counts as a defect.


6. Profile Criteria and Checks​

Click Criteria to choose what the profile checks for. The dialog has two parts.

Built-in checks — Toggle any of these standard checks on or off; one carries a threshold you can set:

CheckWhat it catches
NOT NULL violation (catalog-declared)Empty values in a column the catalog says must be filled.
High null rateA column with more blanks than your set threshold (a null-percentage you choose).
Type drift (declared numeric)Values that don't fit a column declared as numeric.
Duplicate primary-key valuesRepeated values in a primary-key column.
Low cardinality (single value)A column that only ever holds one value.

Custom checks — Add your own checks, each scoped to a column (or, for structural checks, to all columns, a name pattern, or a data type). Available kinds include: not-null, regular-expression match, value-in-a-list, reference-list membership, numeric range, maximum length, and format validators for email, phone, postal code and checksum (Luhn, IBAN, ISBN). There's also an external validation check that calls a configured provider.

In AI mode, a ✦ Suggest with AI button proposes checks based on the sampled data and the catalog — as drafts you accept or edit before they apply.

When you click Apply, the profile re-runs with your criteria. Your criteria are remembered with the session.


7. Running the Profile​

Click Run Profile to profile the sample. The profile rail on the right then shows:

  • A readiness score out of 100 — a single measure of how ready the table is, based on the checks.
  • A narrative summarising the state of the data. A chip marks whether it's an ✦ AI-generated summary or a deterministic one.
  • A Findings tab listing every issue detected.

A dismissible status line at the top of the rail recaps the last run — for example, how many findings were found and the readiness score, and whether AI was actually used.


8. Findings and Promoting Fixes​

Each finding is a card in the Findings tab showing its severity, a description of the issue, and a suggested remediation. In AI mode the explanation is tagged AI; otherwise it's a template.

From a finding you can promote it into a draft fix elsewhere in OnCoor:

Promote toWhat it createsWhere you review it
→ Rule PilotA draft data-quality rule that detects the issue going forward.Rule Pilot (as a pending rule)
→ Cleansing transformA draft cleansing transformation — an expression that corrects the values.Cleansing (as a new draft mapping)

How to promote a fix

  1. Run a profile so the finding you want to act on is showing in the Findings tab.
  2. Decide the mode. In Deterministic mode the draft is built from a template; in ✦ AI mode the button is marked with a ✦ and the AI drafts the actual rule SQL or cleansing expression for you. Switch the toolbar toggle before promoting if you want AI to draft it.
  3. On the finding card, click → Rule Pilot or → Cleansing transform.
  4. A success message confirms the draft was created and gives its artifact number, and the button turns into a green marker (for example, ✓ Rule Pilot #1234) so you can see it's done. The same promotion is also listed on the session report's Actions tab.
  5. Open the draft in its home module (Rule Pilot or Cleansing) to review, adjust and approve it. Nothing is live until you do.

For duplicates, promotion works a little differently — see Duplicate Detection.

PROMOTES CAN BE DISABLED — If the object hasn't been modeled in the platform yet, the promote buttons are switched off. See Modeling in the Platform.

PROMOTIONS ARE ALWAYS DRAFTS — Every promote creates a draft for review in the target module — it never changes live rules or data on its own. AI, when on, only drafts the content; you still approve it.


9. Duplicate Detection​

Click DeDup to look for duplicate records in the sample. Duplicate detection compares rows on a signature — the set of columns you choose in the DeDup on panel (revealed once a sample is loaded).

Pick the signature columns as chips, or use the quick-selects:

Quick-selectWhat it picks
AllEvery selectable column.
NoneClears the selection.
Keys onlyJust the primary-key column(s) — available when the table declares a primary key.
SuggestedThe non-key text columns (the default), which usually make the most meaningful signature.

Results appear in the Duplicate clusters tab, where each cluster can expand to show its member rows. In deterministic mode, matching is textual (near-identical values); in AI mode, it also groups records that mean the same thing.

To act on duplicates, use Add to rule at the top of the tab. This promotes the de-duplication configuration — the match-key columns — as one Cognitive De-Duplication rule for the object. It doesn't promote those specific groups; the groups are re-validated later when the rule runs in Cleansing → Cognitive De-Duplication.


10. External Validation​

If any of your criteria use an external validation check, a Reload validations button appears. Clicking it calls the external provider for the sample's distinct values (for example, verifying addresses or phone numbers), stores the verdicts, and refreshes the profile so the results show up as findings.

THIS IS THE ONLY ACTION THAT SENDS DATA OUT — Reload validations is the single action that sends data to a third-party provider. Running a profile never calls the provider — it only reads verdicts that were already stored. If a provider's usage budget is reached, some values are left unvalidated and the status says so.


11. Modeling in the Platform​

Promoting findings and duplicates requires the object to be modeled in the platform — i.e. to have been through Intake so it has its business attributes and conform-data set up. If it hasn't, an information banner appears and the promote buttons are disabled.

To fix this, pick a Domain layer in the banner (Cleansing by default) and click Model in platform. That runs Intake to populate the object, after which the Rule Pilot, Cleansing transform and De-Duplication promotes all become available.


12. Reports​

A Discovery report brings a table's whole profiling story together in one place — its readiness, criteria, findings, duplicates, the actions taken, and how it's changed over time. There are two ways in, and both open the same report view.

The current session's report — With a session open, click Report in the top-right. It opens that session's report inline, ready to read or download.

The Reports library — From the start screen, click Reports. This lists every table that's been profiled, showing each table's latest session. Filter by Vendor system, then read the row for a quick comparison — Readiness, Findings, Dup clusters, Method (deterministic or AI) and Status — and click View to open the full report.

Inside a report

At the top of a report you can Download PDF, Download Word or Download Excel. If the table has been profiled more than once, a Progress across sessions chart shows how readiness, findings and cluster counts have moved. Below that are tabs:

TabWhat it shows
SummaryHeadline figures — readiness out of 100, average conformance score, finding counts by severity, duplicate clusters and actions — plus a findings-by-severity chart and the key facts (source object, vendor, discovery mode, platform object, sample size, methods, narrative, and when it was generated).
Criteria & MethodExactly what was run — the profile and dedup methods, the dedup signature columns, every built-in check (enabled, threshold, and whether it was default or customised), and any custom checks.
FindingsEvery finding with its severity, kind, column, rows flagged, a conformance Score % (100 = clean), and what it was promoted to.
DeDupCluster totals and the signature columns, plus each cluster's similarity, row count and review state.
ActionsEvery promotion made from this session — the action, its target, and the draft artifact number.
HistoryA timeline across all sessions for the table — who ran each, when, the method, readiness (with the change since the previous run), finding and cluster counts, and status.

REPORTS COME FROM PROFILED TABLES — A table only appears in the Reports library once it's been profiled at least once. The report always reflects the table's latest session.


13. Reference Data and Providers​

Open this hub with Reference & Providers on the start screen. It's where you load the reference lists and manage the external services that your checks validate against. (Profiling doesn't depend on it, which is why it sits on the start screen rather than the session toolbar.) It has two parts: codeset providers and API providers.

Codeset providers — loading a reference list from a file​

A codeset is a static reference list with no live API — for example, a UNSPSC classification set. You load it once from a file, and OnCoor houses it as a reference table you can validate against. Each provider gets its own section showing its default table, how many levels it has, and a status of not loaded, loaded or registered.

Loading a file

  1. In the provider's section, click Load file (or Load / replace file if one is already loaded).
  2. In the dialog, pick the Target connection — the database where the reference table will be created. The dialog also shows the accepted file types (typically .csv, .tsv or .txt).
  3. Click Select file & load and choose your file.
  4. The file uploads with a progress percentage (large files are streamed, so there's no browser size limit), then OnCoor parses and loads it in the background, showing a live row count as it goes.
  5. When it finishes, a message confirms how many codes were loaded.

LOADING REPLACES EXISTING DATA — Loading a file for a provider does a truncate-and-load: it replaces whatever was there before for that provider. The reference table is created on the connection you pick, so validate against a session that profiles that same connection.

After loading, each dataset is listed with its table name, code count, level breakdown, source file and size, status and last-updated time. From here you can:

  • Register & set up — Turns the loaded table into a governed Reference Config: it automatically creates the columns, Conform Data, and the CODE lookup join, then makes the codeset appear in the REFERENCE check's Reference Config picker. Re-run it any time as Re-register to refresh. You give it a Reference name and a Domain layer when registering.
  • Remove (✕) — Drops the dataset's record. The underlying table is left intact.

Using a loaded codeset — Once registered, add a REFERENCE custom check in Profile Criteria, point its Reference Config at your registered codeset, and set the column to CODE. The profile then flags any value that isn't in the list.

API providers​

The lower section lists the external validation providers (for email, phone, postal, address and similar checks) available to this data domain, with a chip showing whether external validation is switched ON or OFF overall. For each provider you can see its type, whether it's enabled, its per-run budget, and whether it uses an API key or is keyless.

This view is read-only — a provider's enablement, budget and credentials are governed centrally in Admin → Discovery Providers, per data domain. This hub simply reflects the current domain's settings so you know what's available when you build an external validation check.

A PROVIDER WITH DATA BUT NO ACCESS — If a codeset was loaded before but its provider is no longer enabled for the domain, it still shows here (flagged "not available in this domain") so its data isn't lost — re-enable it in Admin → Discovery Providers to load or replace it.


14. Quick Reference​

A fast lookup for the most common actions.

I want to…Do this
Profile a catalogued tableStart screen → From catalog → pick object → Start session
Profile a live table with no catalogStart screen → Direct from source → pick table → Start session
Reopen earlier workStart screen → Resume a session list
See real rowsOpen a session → Load sample
Add AI explanations, matching and draftsToolbar → set mode to ✦ AI
Choose what the profile checksCriteria → toggle built-ins / add custom checks → Apply
Score the table and list issuesRun Profile
Draft a fix from a findingOn the finding card → → Rule Pilot or → Cleansing transform
Find duplicate recordsPick DeDup on columns → DeDup
Draft a de-duplication ruleDuplicate clusters tab → Add to rule
Validate against an external serviceAdd an external validation check → Reload validations
Enable promotes for a new objectBanner → pick a domain layer → Model in platform
Get the report as a fileReport → Excel / PDF / Word
Compare a table's sessions over timeReports → open a table → History tab
Load a reference list from a fileReference & Providers → provider → Load file
Make a loaded codeset usable in checksReference & Providers → Register & set up

Source: OnCoor Discovery Module product documentation, written for end users. For the latest screens and options, always refer to the in-app interface.