Nothing Happened to the Spelling
Names are evidence of identity, not identity itself. Naming systems preserve crossovers, spelling variation and publication thresholds that make a clean join look more certain than it is.
- You are seeing
-
- An exact name match triggers a customer merge without a second independent field
- A historical employee series loses people after spelling standardization
- A rare name disappears from a vendor file and is treated as a departure
- A surname field is used as a stable lineage key in a regenerated naming system
- A report treats one spelling's rank as the prevalence of every variant
- The mechanism
- Given-name and surname records both show that one string can cross categories, change spelling, and fall below reporting thresholds, so matching strings is weaker than resolving people or entities.
- The older apparatus
- Onomastics records origin, spelling, geographic distribution, and administrative classification separately, rather than treating the printed name as a unique key.
- The false friend
- A duplicated record of the same entity also yields two similar names, but it shares a stable source identifier or an impossible duplicate transaction history.
- The discriminating test
- Check whether the match survives a second independent field such as date, address lineage, jurisdiction, or relationship; name-only matches that split under a second field are collisions.
- On your own data
- Score entity links by independent fields, retain unresolved candidates, and report results under strict and permissive linkage rules rather than emitting one forced merge.
In 1984 Splash gave the United States a reason to hear Madison as a given name.
Before 1985 it was essentially unused as one. It had been an English surname, meaning son of Matthew or son of Maud. By 2001 it ranked second among American girls’ names.
Nothing had happened to the spelling. The string was identical throughout. The population underneath it had changed.
That is the first error a clean join makes. It treats the printed field as the thing being identified, when the field is an observation made by a naming system at a particular moment.
Two populations under one spelling #
Madison is not a curiosity made interesting by a film. It holds several distinct histories inside a single word.
Before its rise among girls it had a minor place in the American boys’ top thousand until about 1952. It dropped out, then re-entered from 1987 to 1999, never approaching the scale of the girls’ rise. Those are effectively two separate populations occupying one spelling, and both are live.
Ashley does the same thing over a longer interval. An English place-derived surname, ash tree meadow, used as a boys’ given name from the 1860s, still ranked 33rd for boys in England and Wales in 1994. In the United States it had shifted to majority-female use from the 1940s and accelerated after 1982.
A matcher handed Madison or Ashley is not deciding what either name means. It is deciding whether this particular occurrence belongs to a surname population, a boys’ given-name population, a girls’ given-name population, or a person who merely shares the letters. The word’s origin does not settle that. Neither does the exactness of the spelling.
The convenient story is that a name begins as one thing and later becomes another. The records are less tidy. The old use does not leave when the new use arrives; it stays, attached to the same letters, filed in the same column.
Counting rules change the size of a name #
The Social Security Administration’s baby-name series is a complete count of card applications for every birth year from 1880 onward. Its fields are name, year of birth, sex and count. There is no region, no parental characteristic, no semantic origin, and no decision that two spellings are one name.
Rank is derived by sorting the count. Identity is not among the recorded fields.
That matters most where the spelling line looks innocuous. In 2006 the SSA ranked Mohammad 589th, Mohammed 633rd and Muhammad 639th. Finland’s statistics aggregate the variants instead, recording some 15,700 people carrying a form of Muhammad, with Mohamed making up well over a third of the total. Neither system has discovered a different population. Each has adopted a different tabulation rule.
The same distinction governs the tail. A name below a publication threshold can be legally registered and in daily use while absent from the published series. That absence records the line in the report. A vendor row that disappears after a rare-name filter changes has not, on that evidence, gone anywhere.
Surname files make the line explicit. The 2010 US Census surname release covers names occurring a hundred or more times, some 162,000 of them, and reports that fewer than two thousand surnames account for half the population. A file built at that threshold measures the concentrated half accurately and is structurally silent about most of the thin tail.
A surname is not always a lineage #
Icelandic patronymics make the sharper point, because they look exactly like the field a database expects.
An Icelandic -son or -dóttir name is routinely filed as a surname, and it is not inherited unchanged. Ólafur Jónsson’s children are not Jónsson. Their patronymic is built from Ólafur, not from Jón.
The 1925 law preserved a small existing population of genuine fixed family names, a few hundred holders at most, without making fixed names the ordinary rule. So the current system holds a legacy fixed-surname minority beside a patronymic and matronymic majority, in one field, under one label.
Calling both groups surnames is not a terminological slip. It changes the statistic. A frequency count of Icelandic surnames is a count of inherited lineages for one group and a count of given-name popularity one generation back for the other. Standardise the field and the series has joined two measurements that were never measuring the same thing.
Stability can also be manufactured outright. Turkey’s 1934 Surname Law required a Turkish-language hereditary surname within the year and prohibited the Armenian, Slavic, Greek and Persian endings that had marked origin. The law’s characterisation remains disputed. Its effect on the string does not: a marker can stop appearing because a naming rule demanded a different one.
Where the transfer stops #
Names are good evidence because they carry history — a crossover, a spelling convention, a reporting threshold, a language, a classification imposed by a registrar. They are poor unique keys for precisely the same reason.
People carry that ambiguity for life. A name can move between surname and given-name use, pick up a variant spelling, fall below a publication line, or be changed by administrative decree, and every affected record still prints something entirely valid.
Account systems are built the other way. An issuer identifier is designed to suppress history rather than preserve it, which is what makes it a usable key. That advantage is real and it is not permanent: mergers and vendor remaps sever identifiers too, and the replacement starts accumulating a history of its own from the day it is issued.
An exact string match is the cheapest evidence a join will accept, and it is the evidence a name was least built to supply.
Madison is spelled the same way in 1950 and in 2001. Everything that changed was something the spelling was never recording.
Whatever decides that a merge combined one entity rather than two will be the other field. The one that had no particular reason to agree, and agreed anyway.
Questions
Can an exact name match identify the same customer?
No. An exact name match is one piece of identity evidence, not a unique identifier. Madison is an English patronymic surname that was essentially unused as a US first name before 1985, then ranked second among girls' names by 2001. The same spelling therefore belongs to different naming categories and many people. A date, address lineage, jurisdiction, relationship, or stable source identifier is needed to distinguish an entity link from a collision.
Why does spelling standardization make people disappear from a series?
Because the tabulation may treat each spelling as a separate record even when the spellings refer to one name lineage. The US Social Security Administration ranked Mohammad 589th, Mohammed 633rd, and Muhammad 639th in 2006. Finland instead aggregates variants and reports 15,723 people carrying some form of Muhammad. A standardization can therefore split, combine, or erase continuity unless the variant rule is recorded with the series.
Does a name missing from a published file mean the person has left it?
No. Absence from a published name file can mean that the name falls below that file's reporting threshold. The US Census Bureau's 2010 surname release includes only surnames occurring 100 or more times, while its published file contains 162,253 distinct surnames. A rare surname can remain in use while disappearing from the release, so a missing row is not evidence of a departure without an independent record.
Can a surname be used as a stable family-lineage key everywhere?
No. Icelandic -son and -dóttir names are regenerated from a parent's given name rather than inherited unchanged. Ólafur Jónsson's children take a patronymic built from Ólafur, not Jón. A frequency count filed as Icelandic surnames therefore measures the popularity of given names one generation back, not a fixed lineage. A linkage system must identify the naming regime before using the field as a persistent key.
When does a duplicate-looking record show duplication rather than a name collision?
A duplicate record has evidence that both rows refer to the same underlying entity beyond the name itself. That evidence can be a stable source identifier, a consistent address lineage, or a transaction history that cannot belong to two entities. A collision is the opposite condition: the name match breaks when an independent field is checked. The distinction determines whether records can be safely collapsed or must remain competing candidates.
Sources
- "Ashley (given name)," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Baby Names from Social Security Card Applications - National Data," Data.gov catalog (primary, 2026-08-13)
- "Black names," Wikipedia (Fryer and Levitt 2004; University of Chicago/UC Berkeley callback study) · Wikimedia Foundation (tertiary, 2026-08-13)
- "Catálogo alfabético de apellidos," Wikipedia (Philippines, 1849 Clavería decree) · Wikimedia Foundation (tertiary, 2026-08-13)
- "Icelandic name," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Korean name," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Madison (given name)," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Muhammad (name)," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Naming law," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Nevaeh," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Surname Law," Wikipedia (Turkey, 1934) · Wikimedia Foundation (tertiary, 2026-08-13)
- "Surname," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Unisex name," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Vietnamese name," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- "Welsh surname," Wikipedia · Wikimedia Foundation (tertiary, 2026-08-13)
- Office for National Statistics, "Baby names for England and Wales: 2023" (primary, 2026-08-13)
- US Census Bureau, "2010 Census Surnames" · U.S. Census Bureau (primary, 2026-08-13)