Cognitive De-Duplication
User Guide — Finding and Resolving Duplicate Records in an Enterprise Object
| Last updated: August 2026
Contents
- Overview
- Key Concept: One Rule, Four Steps, Match Groups
- Turning On Cognitive De-Duplication
- Finding Your Way Around
- Step 1 — Configure Rules
- Step 2 — Process Setup
- Step 3 — Match Group Dashboard
- Step 4 — Review & Submit
- Quick Reference
1. Overview
Cognitive De-Duplication finds records that refer to the same real-world thing — the same customer, material, or supplier recorded more than once — and helps you resolve them. It can match on exact values, tolerate typos, or use AI to spot matches that don't look alike on the surface but mean the same thing.
You set it up once per Enterprise Object as a rule: you tell it which columns identify a duplicate, how strict to be, and how to match; then you run it, review the groups of records it found, approve the ones that really are duplicates, and submit the result. The outcome is a reference table — a crosswalk that links each duplicate to its survivor — which downstream pipelines use to consolidate the data.
What you can do here
- Choose the columns that identify a duplicate — the "match keys."
- Pick how strict matching is — exact, typo-tolerant, or AI-semantic.
- Run the matching across the object's data and watch its progress.
- Review the match groups it found, with a confidence score for each.
- Approve or reject each group — individually or in bulk.
- Submit the approved matches into a reference table for downstream use.
- Schedule the matching to re-run on a recurring basis.
WHERE TO FIND IT — Cognitive De-Duplication is a tab on the Enterprise Object (shown as Cognitive DeDup). Open the object, choose the layer, and select the tab. There is one de-duplication rule per object.
NOT THE SAME AS RULE PILOT'S DUP CHECK — Rule Pilot has a "Dup Check" that flags repeated values in a column as a quality rule. Cognitive De-Duplication is the fuller capability: it matches whole records (exactly, fuzzily, or with AI), lets you review and approve the matches, and produces a reference table to consolidate them. Use Dup Check for a simple quality flag; use Cognitive De-Duplication to actually resolve duplicates.
2. Key Concept: One Rule, Four Steps, Match Groups
One rule per object. The whole feature is a single de-duplication rule attached to the Enterprise Object. The first time you open the tab on a new object, you create the rule; after that you edit and re-run the same one.
Four steps. You work through a wizard of four steps, shown as a row across the top. Later steps unlock as earlier ones are satisfied.
| # | Step | What you do |
|---|---|---|
| 1 | Configure Rules | Choose the tables and columns that identify a duplicate. |
| 2 | Process Setup | Choose how matching works, then run it. |
| 3 | Match Group Dashboard | Review the groups of matching records and approve or reject them. |
| 4 | Review & Submit | Confirm the approved matches and write them to a reference table. |
Match groups. When the rule runs, it clusters records that appear to be the same into match groups. Each group gets a match score (how confident the engine is) and a category (Exact, High, or Medium confidence). Reviewing means deciding, group by group, whether it's a real duplicate — Approve it — or not — Reject it. The survivor. Within each match group, one record is designated the survivor — the record that's kept and that the others are linked to (in the reference table, every duplicate's key points to the survivor's key). OnCoor picks a survivor for each group automatically when the rule runs; you can override that choice during review (see The group grid).
STEPS UNLOCK IN ORDER — Process Setup needs at least one match-key column from Configure Rules. The Match Group Dashboard and Review & Submit only open once the rule has been run at least once. Disabled tabs show a tooltip telling you what's needed.
3. Turning On Cognitive De-Duplication
Cognitive De-Duplication is off by default. If the Cognitive DeDup tab is greyed out (or the run steps never unlock), the feature hasn't been switched on for that layer yet. Turning it on is a one-time setup in the Enterprise Object configuration, and it's done per layer — enable it on whichever layer you want to de-duplicate in (for example, the Cleansing layer).
3.1 Before you start
Two things need to be in place:
- Edit rights on the configuration. Without them the switches are read-only.
- The domain must allow capabilities. A domain-level setting, Allow Domain Capabilities, has to be on. While it's off, only Extraction can be enabled and the De-duplication switches stay greyed out (the screen shows a note saying so). That setting lives on the Domain tab of the configuration — an administrator turns it on first.
3.2 Enabling it for a layer
- Go to Enterprise Object Configuration → Domain Layer.
- Select the layer you want to de-duplicate in.
- Find the Capabilities section. It lists five capability pairs — Profiler, De-duplication, Extraction, Validation, and Data Approval — each with two switches: Capable and Enabled.
- On De-duplication (hint: "Remove duplicate rows during conform"), switch on Capable first — this says the layer is allowed to de-duplicate.
- Then switch on Enabled — this turns it on right now. (Enabled stays locked until Capable is on.)
- The change saves to the layer.
Reopen the Enterprise Object and the Cognitive DeDup tab is now active for that layer.
CAPABLE VS ENABLED — "Capable" means the layer is allowed to de-duplicate; "Enabled" means it actually does. You need both on. Because turning Capable off also turns Enabled off, you can pause the feature later — without losing its rule and settings — by leaving Capable on and switching only Enabled off.
IT'S THE SAME SWITCH FOR THE OTHER TABS — The Capabilities section is also where the object's other tabs are turned on: Profiler powers Rule Pilot, Extraction powers Data Gen, plus Validation and Data Approval. Enabling De-duplication on one layer doesn't turn it on elsewhere — set it on each layer where you need it.
4. Finding Your Way Around
Open the Enterprise Object, choose the layer, and select the Cognitive DeDup tab. The page shows the rule's name at the top and the four step-tabs beneath it. If the object has no rule yet, you'll see a + Create Rule button — click it to start.
The rest of this guide follows the four steps in order.
5. Step 1 — Configure Rules
Configure Rules is where you tell the engine what makes two records a duplicate. The screen has three panels side by side: the object's tables on the left, the selected table's columns in the middle, and a Selected Columns summary on the right.
5.1 Naming the rule
At the top, give the rule a clear Rule Name so it's easy to recognise later.
5.2 Choosing match keys and context columns
Pick a table on the left, then click its columns in the middle to give each a role. Clicking a column cycles it through three states:
| State | Colour | Meaning |
|---|---|---|
| (none) | White | The column isn't used. |
| Match key | Cyan | Drives the matching — these are the columns the engine compares to decide if two records are duplicates. |
| Context | Purple | Not used for matching, but shown alongside the records when you review, to help you judge a group. |
Click once to make a column a match key, again to switch it to context, and a third time to remove it. A rule needs at least one match key — context columns alone can't find duplicates.
You can configure columns across several tables; a ✓ marks each table that already has columns set.
5.3 The Selected Columns summary
The right-hand panel lists everything you've configured, grouped by table, with a count of keys and context columns. From here you can jump to a table, remove a single column, clear one table's columns, or clear the whole rule.
When at least one match key is set, Initiate De-Duplication → takes you to Process Setup.
PICK THE RIGHT MATCH KEYS — Match keys are the heart of the rule: too few and unrelated records get grouped together; too many and real duplicates slip through. Start with the columns a person would use to recognise a duplicate (name, code, key identifiers), and add context columns for anything that just helps you eyeball a group during review.
6. Step 2 — Process Setup
Process Setup is where you choose how the matching behaves and then run it.
6.1 Match method
The Match Method decides how records are compared — this is what actually controls the matching:
| Method | How it matches | Speed | Best for |
|---|---|---|---|
| Exact | A direct, case-insensitive, whitespace-tidied comparison of the match-key columns. | Fastest — typically seconds. | Clean, well-standardised data. |
| Fuzzy | Typo-tolerant similarity on each column, using the threshold below as the cutoff. | Moderate — grows with row and column counts. | Short text fields with typos and abbreviations. |
| Cognitive | AI that understands each record's meaning and groups by semantic similarity. | Slowest — depends on data volume and the AI provider. | Genuinely different-looking values that mean the same thing, across several noisy columns. |
HOW LONG A RUN TAKES — There's no fixed time — it scales with how many rows and columns the rule compares, and Cognitive also depends on your AI provider's speed. Exact is effectively instant on typical volumes; Fuzzy and especially Cognitive take longer on large tables. The run happens in the background with a live progress bar, so you can start it and come back.
6.2 Process objective
The Process Objective records your intent for what happens to the matches downstream. It's descriptive metadata — it doesn't change how matching runs (that's the Match Method).
| Objective | Meaning |
|---|---|
| Consolidation of Source Records | Merge matched records into a single golden record using survivorship rules. |
| Deduplication | Keep one survivor per group and mark the others inactive in the source. |
| Record Linkage | Build a permanent crosswalk between matches; the source data isn't modified, and downstream pipelines join to the reference table. |
On the Cleansing layer, this defaults to Record Linkage, because its reference-table output is what consolidates source records downstream.
6.3 Match threshold
For Fuzzy and Cognitive methods, a Match Threshold slider (0–100%) sets how similar two records must be to be grouped — only pairs at or above the threshold are matched. Exact matching is all-or-nothing, so the threshold doesn't apply to it.
6.4 Blocking keys
For the Cognitive method, a Blocking Keys panel lets you narrow which records are compared to each other, so the engine doesn't compare every record against every other one. This keeps large runs fast and focused.
6.5 Schedule
A Schedule panel lets you set the matching to re-run automatically on a recurring basis, so duplicates are caught as new data arrives without someone running it by hand.
6.6 Running and monitoring
Click → Execute Process to start a run. Matching runs in the background, so you can leave the page or close the tab — it keeps going. A banner shows live progress: the percentage complete, a status message, and how long it's been running, with a Cancel run button. When it finishes, the banner reports the outcome:
| Outcome | Meaning |
|---|---|
| Completed | The run finished; the match groups are ready to review. |
| Failed | Something went wrong; the message explains what. |
| Cancelled | You stopped the run before it finished. |
A completed run unlocks the Match Group Dashboard and Review & Submit steps.
6.7 Using the result downstream (Cleansing layer)
On the Cleansing layer, the rule produces a reference table for each master table — a crosswalk linking duplicates to their survivor. A guidance panel on this step shows the exact reference tables the rule will create and walks through wiring each dependent transaction table to them via Mapping Manager → Reference Join, so the deduplication actually flows through to the data that depends on it.
A RUN IS SAFE TO LEAVE — Because matching runs in the background, you don't have to sit and watch it. Start it, leave, and come back — the banner (and the unlocked review steps) will be waiting when it's done.
7. Step 3 — Match Group Dashboard
The dashboard is where you review what the run found and decide what's a real duplicate. Records the engine believes are the same are clustered into match groups, one row per group.
7.1 Summary and scope
KPI cards at the top show how many groups and records fall into each confidence category. A source-table strip lets you scope the whole view to one table (or Cross-table, for groups that span tables), and the KPI cards re-scope to match.
7.2 Filters
Two filters narrow the list:
- Status — All, Open (not yet decided), Approved, Rejected, or Mixed.
- Category — Exact, High, or Medium confidence.
7.3 The group grid
Each row is a match group:
| Column | What it shows |
|---|---|
| Match Group # | The group's identifier. |
| Records | How many records are in the group. |
| Match Score % | The engine's confidence that they're the same. |
| Category | Exact, High, or Medium. |
| Group Decision | Whether the group is Open, Approved, or Rejected. |
| Record State | The state of the records within the group. |
| Actions | Approve, Reject, or Reset the group's decision. |
Expanding a group shows its member records with the context columns you configured, so you can confirm they really match before deciding. Each group also has a survivor — the record the others will be linked to, chosen automatically by the engine and marked in the panel. If the wrong record was picked, use Set as survivor on the record you want to keep (or Reset to return to the engine's choice); OnCoor re-points the group's pairs to the new survivor.
7.4 Approving and rejecting
Decide each group with its row Actions — Approve a real duplicate, Reject a false match, or Reset to clear a decision. To move quickly, use Approve all open or Reject all open to decide every undecided group at once (already-decided groups are left untouched), or select several groups and act on them together.
7.5 Exporting
Export downloads every group and its member records to Excel — one block per group, mirroring the dashboard — for offline review or sign-off.
ONLY APPROVED GROUPS ARE WRITTEN — Rejecting a group leaves those records untouched; only groups you approve are carried into the reference table at Submit. Groups left Open are not submitted, so work through them before moving on.
8. Step 4 — Review & Submit
The final step is a read-only confirmation of everything you approved, grouped by source table, with a single Submit at the bottom. It doesn't re-open the decisions — those are made on the dashboard — it confirms exactly what will be written.
Each source table's tab shows the approved matches as From → To pairs (each duplicate mapped to its survivor), along with who created them and when. A per-table count separates direct matches (pairs the engine actually found) from propagated ones (related child records that follow along because of a foreign-key relationship), so you can check the result landed where you expect.
Submit commits the full approved set into the per-source reference tables (named dedup.<source>_DEDUP_REF). Once written, those tables are what downstream pipelines join to in order to consolidate the duplicates. A confirmation dialog shows the direct/propagated split before you commit, and a summary afterward reports how many reference tables and pairs were written.
SUBMIT IS THE COMMIT — Approving groups on the dashboard records your decisions; Submit is what actually writes them into the reference tables downstream data depends on. Make sure you've reviewed the open groups before submitting.
9. Quick Reference
A fast lookup for the most common actions.
| I want to… | Do this |
|---|---|
| Open Cognitive De-Duplication | Enterprise Object → choose the layer → Cognitive DeDup tab |
| Start a rule on a new object | + Create Rule |
| Say what makes a duplicate | Configure Rules → click columns to set Match keys |
| Add columns just to help review | Click a column twice to make it Context |
| Choose how strict matching is | Process Setup → Match Method (Exact / Fuzzy / Cognitive) |
| Set the similarity cut-off | Match Threshold slider (Fuzzy / Cognitive only) |
| Speed up a large AI run | Set Blocking Keys (Cognitive) |
| Re-run automatically | Set a Schedule |
| Run the matching | → Execute Process |
| Stop a run | Cancel run in the progress banner |
| Review what was found | Match Group Dashboard |
| Focus on undecided groups | Set the Status filter to Open |
| Accept / reject a duplicate | Approve / Reject on the group's row |
| Decide everything at once | Approve all open / Reject all open |
| Get the groups into Excel | Export |
| Commit the approved matches | Review & Submit → Submit |
| Use the result downstream | Wire each table via Mapping Manager → Reference Join (Cleansing layer) |
Source: OnCoor Cognitive De-Duplication product documentation, written for end users. For the latest screens and options, always refer to the in-app interface.