Record Linkage Scorer

Score possible cross-dataset record pairs with explicit blocking and field evidence. Nothing is uploaded or merged.

Left dataset

Right dataset

Candidate generation and scoring

Methods: exact, normalized, token-set Jaccard, Unicode edit, decimal numeric, and ISO YYYY-MM-DD date. Similarity thresholds are minimums; numeric/date thresholds are maximum distance (number or days). Field thresholds annotate evidence; the overall decision uses available-field similarities with positive weights renormalized after missing or invalid values are omitted.

Metrics use supplied labels only. Precision ignores unlabeled predictions; recall treats supplied positive labels removed by blocking as false negatives.

Limits per dataset: 3 MiB, 25,000 rows, 200 columns. Limits per run: 25 rules, 100,000 labels, 100,000 candidate pairs, 500,000 field comparisons, and 20 million estimated comparison operations; values over 2,000 characters are marked unavailable. Preview: 50/page. Missing blocking values create no candidate and are reported. Duplicate rows remain separate by row number.

Scored candidates

Candidate pair scores and per-field evidence
Left rowRight rowScoreDecisionPer-field evidence

Findings