Lab

CRM dedupe & standardiser

Drop in a messy contact export, get back one clean record per person, with every merge explained and nothing sent to a server.

Adjust inputs ↓

–

–Records built from 2+ rows
–Pairs in review
–Cells standardised

Pipeline · 5 steps

Same file in, same records out

    Candidate pairs

    Where the thresholds cut

    Kept apart Review Auto-merge
    SignalWeightHow it is measured

    Check each merge before you export

    Golden recordRowsWeakest linkSignalsDetail
    How it works
    Method
    Rule-based standardisation, blocked record linkage (only records that share a key are compared), weighted similarity, union-find clustering, rule-based survivorship (rules pick which row supplies each field of the merged record).
    Built from
    crm-refine, my open-source CLI (Python, no dependencies), ported to browser JavaScript.
    Data
    Synthetic, generated from a seeded model. Fictional people and companies. Not client data.
    Privacy
    Your CSV is read in this tab. No network calls.

    1. Standardise

    Each cell is cleaned before anything is compared. Names are trimmed and cased (Mc, Mac, O', hyphens, and particles like von or van). A full-name column is split, including "Last, First". Emails are lower-cased with plus-tags removed. For Gmail and Googlemail only, dots in the local part are ignored when matching, because Google ignores them too. Phones become E.164 (+4930…) using the row's country, or the default country if the row has none. Legal forms like GmbH & Co. KG, AG, Inc. and Ltd are stripped from company names. Countries become ISO-2 codes. Job titles get a seniority and a function from ordered rules. Lifecycle stages map to one ladder and dates become YYYY-MM-DD. US rows read 03/04 as March 4; others as 3 April.

    2. Block

    Comparing every row with every other row does not scale. A record only meets records that share a key: the same normalised email, the same E.164 phone, the same company plus the same phonetic name, or the same phonetic name alone. Phonetic codes use Kölner Phonetik (Cologne phonetics), so Müller, Mueller and Muller share a code. First names go through a nickname table first (Bob to Robert, Kate to Katharina). The two name codes are sorted, so swapped first and last names still meet. Blocks with more than 40 records are skipped to keep the work bounded.

    3. Score

    Each candidate pair gets a score from 0 to 1. It is the weighted average of the signals both records have; missing fields are left out instead of counting against the pair. Names use Jaro-Winkler on folded, nickname-resolved names, and are also compared swapped. Company uses token Jaccard on the crm-refine fingerprint. A different email counts at half weight, because people keep work and personal addresses. Two guards push a pair down into review: when only the name matches, and when the first names clearly differ, which usually means a shared inbox like j.carter@.

    4. Cluster

    Pairs at or above the auto-merge line are joined with union-find, the same structure crm-refine uses, so A–B and B–C put A, B and C in one record. Pairs in the review band are listed for a person to accept or reject. Decisions apply at once and are kept until you load new data.

    5. Survive

    Each cluster becomes one golden record. Fields that change when people move jobs (email, phone, company, title, country) come from the most recently active row. Names come from the most complete value, so Katharina beats Kate. The highest lifecycle stage is kept and the latest activity date wins. Columns you did not map come from the most complete row, filled from a duplicate when blank, as in crm-refine. Every field records which row it came from and why.

    What came from crm-refine

    Ported as-is: the company fingerprint and legal-suffix list, the free-mail domain list, name casing, company casing, the phone rules (country dial codes, trunk zero), the country table, union-find, master selection by completeness then recency, blank-field fill from duplicates, the auto versus review split, and pairwise precision and recall. crm-refine matches companies with sorted-neighbourhood edit distance. Person-level matching is not in the CLI, so this page uses the standard equivalents: Jaro-Winkler, a nickname table, Cologne phonetics and weighted scoring.

    In production

    The same pipeline runs as the dependency-free crm-refine CLI, or as a scheduled job against a HubSpot or Salesforce export. It writes a clean file and a merge log. A person reviews the log before anything is written back, because merges are hard to undo in most CRMs. An LLM can classify messy job titles at scale. The matching stays deterministic so every merge can be audited and replayed.

    Limits

    • No external enrichment. It only knows what is in the file.
    • The nickname table is partial and tuned for English and German names.
    • Phonetic keys are language-specific. Cologne phonetics fits German spelling better than English or Spanish.
    • Weights are set by hand, not learned. With labelled pairs you would fit them (Fellegi-Sunter or a small classifier).
    • Same-name colleagues with no email or phone look identical to the model. They land in review at best.
    • The score against ground truth only covers the synthetic sample. On the default sample at default thresholds it reads about 1.00 precision and 0.97 recall.