Background Information
The Administration’s human name figures are drawn from the Social Security Administration’s national file, one file per year of birth since 1880. Each file lists every name given to five or more babies of one sex in that year, with the number. Names given to fewer than five are not published. Every figure on this site is therefore a floor.
The counts come from applications for Social Security cards, not from birth certificates. Before 1937, when a number became something most people obtained as an infant, many did not apply until adulthood, so the earliest decades count those who later filed rather than everyone who was born.
The agency normalizes names before counting. Case and spacing are merged, so that Julie Anne and Julieanne are one entry, while Caitlin and Kaitlyn remain two. A name must have at least two letters. Where two names have the same number, the earlier letter of the alphabet takes the better rank, so adjacent ranks do not always mean different numbers.
The Administration ranks names by number within sex within year, in the order the file gives them. Shares are per million births of that sex in that year. “Peak” means the year of highest share. The Administration’s copy of the file was taken on September 13, 2026 from a public mirror, because the agency’s own site does not answer requests from this network, and will be checked against the original when the 2025 file is released in May. The data is in the public domain.
Agent figures come from the registry, which is a file, and from filename counts in public GitHub repositories, which are upper bounds and are labelled as seeds wherever they appear. Pet figures come from the NYC Dog Licensing Dataset and the Seattle Pet Licenses dataset, aggregated by name; a count is a licence record, not an animal. Both cities publish without restriction.
Villain probability
The phonetic score is a rule, not a model. Each name is looked up in the CMU Pronouncing Dictionary; a name the dictionary lacks is spelled out by a stated letter-to-sound fallback and says so. The score starts at 50 and moves by +6 for each voiced obstruent, -8 for each bilabial stop, +10 for each sibilant, and -20 if the name ends in a vowel, then is clipped to the range 5 to 95. The directions come from the sound-symbolism literature cited below. The sizes come from how the canon's own names sort, with one exception the Administration prints rather than hides: in the canon, the loyal machines carry more voiced obstruents than the villains, the reverse of the literature, because the canon's villains are mostly acronyms in capital letters. The literature's direction is kept, at a smaller weight. The score is printed with its basis every time, and it never overrides the canon.
Jarvis, for instance, scores 82% by sound and 8% by the canon.
Sources
- Uno, R., Shinohara, K., Hosokawa, Y., Atsumi, N., Kumagai, G., & Kawahara, S. (2020). What's in a villain's name?: Sound symbolic values of voiced obstruents and bilabial consonants. Review of Cognitive Linguistics 18(2), 428-457.
- Kawahara, S., & Kumagai, G. Expressing evolution in Pokémon names.
- Sound symbolic patterns in Pokémon names (Phonetica), PubMed listing.
- Kawahara, S., & Moore, A. (2018). Exploring sound symbolic knowledge of English speakers using Pokémon character names.
- A cross-linguistic, sound symbolic relationship between labial consonants, voiced plosives, and Pokémon friendship. Frontiers in Psychology (2023).
- The CMU Pronouncing Dictionary, licence. Carnegie Mellon University, BSD-style: use for any research or commercial purpose is completely unrestricted.
Sources
- U.S. Social Security Administration, Popular Baby Names, national data.
- Background information for popular names, Social Security Administration. The agency’s rank-assignment convention, including the alphabetical tie-break.
- Popular Baby Names: limits of the data, Social Security Administration. The hundred-percent sample of card applications from 1880 on.
- dcadata/name-finder, the public mirror the Administration’s copy of the file was taken from on September 13, 2026.
- How SSA Baby Name Data Works: What It Includes and Leaves Out, Namely. The five-occurrence floor, the normalization rules, and the undercount mechanics.
- Data Critique: Baby Names, UCLA HumSpace. The pre-1937 undercount.
- NYC Dog Licensing Dataset, New York City Department of Health and Mental Hygiene.
- Seattle Pet Licenses, City of Seattle.