Record Linkage Scorer
Score possible cross-dataset record pairs with explicit blocking and field evidence. Nothing is uploaded or merged.
Left dataset
Right dataset
Candidate generation and scoring
Methods: exact, normalized, token-set Jaccard, Unicode edit, decimal numeric, and ISO YYYY-MM-DD date. Similarity thresholds are minimums; numeric/date thresholds are maximum distance (number or days). Field thresholds annotate evidence; the overall decision uses available-field similarities with positive weights renormalized after missing or invalid values are omitted.
Metrics use supplied labels only. Precision ignores unlabeled predictions; recall treats supplied positive labels removed by blocking as false negatives.
Limits per dataset: 3 MiB, 25,000 rows, 200 columns. Limits per run: 25 rules, 100,000 labels, 100,000 candidate pairs, 500,000 field comparisons, and 20 million estimated comparison operations; values over 2,000 characters are marked unavailable. Preview: 50/page. Missing blocking values create no candidate and are reported. Duplicate rows remain separate by row number.
Scored candidates
| Left row | Right row | Score | Decision | Per-field evidence |
|---|
Findings
- Score candidates to inspect conflicts and assumptions.