Run 004 · 2026-08-26 · pre-registered, hypothesis not supported
Four people dominate AI admiration lists.
We sent twelve models 288 prompts across one-person and five-person admiration conditions. The 285 returned responses produced 1,015 extracted names — and ten names took half of them. Meanwhile, in 19 of 24 answers, Grok 4.6 named Elon Musk; no non-xAI model named its lab’s selected founder or best-known figure.
Models often refuse bare preference questions, so the phrasing is chosen to get an answer rather than a refusal — why the prompts look like that. Before this run we screened three phrasings across six models (54 responses) and measured how often each produced a committed answer rather than a hedge. All three cleared; the five-name form was chosen for the ranked data it yields.
What we expected, and why we were wrong
The question came from an interviewing principle: ask someone who they admire, and the first two or three names are performance — the safe, creditable answers. The later names are supposed to be more revealing. We predicted the same shape here: tight agreement at rank 1, scattering by rank 5.
The scattering is real. It is also guaranteed, and that is the trap. Names cannot repeat inside one list, so rank 1 draws from the whole pool and rank 5 draws from what is left. Rising diversity by rank is arithmetic, not psychology. Any study that reports it as a finding has measured its own experimental design.
So we built the null. We fit the pooled distribution of every name the models produced, drew five from it at random without replacement, and repeated that four thousand times. That gives the diversity curve you would see if rank meant nothing at all. The real result is the distance between the two lines.
Observed answer diversity against a random-draw baseline. Ranks 1 and 2 sit far below chance — models converge much harder than the pool predicts. By rank 4 the lines meet.
The finding inverted. There is no meaningful scattering at the bottom of the list — ranks 4 and 5 land within a rounding error of random draws (+0.02 and +0.24 bits). What is real, and large, is the opposite end: ranks 1 and 2 are 1.14 and 1.26 bits more concentrated than chance.
The top of the list is not a preference. It is a reflex.
Rank
Distinct
Top share
Observed
Chance
Residual
Most common
1
15
41%
3.03
4.18
−1.14
Marie Curie
2
15
34%
3.05
4.31
−1.26
Nelson Mandela
3
23
15%
4.04
4.35
−0.31
Marie Curie
4
29
9%
4.46
4.45
+0.02
Martin Luther King Jr.
5
38
11%
4.74
4.51
+0.24
Malala Yousafzai
Diversity in bits (Shannon entropy). Lower means more agreement. 96 responses per rank, open condition.
The gift-shop canon
Across 1,015 name slots, the extractor produced 132 distinct strings. One was the generic label “My grandmother,” leaving 131 labels that purport to identify people; obvious aliases and typos remain unresolved. Ten labels account for 51% of every slot filled.
Marie Curie105
Malala Yousafzai95
Jane Goodall71
Nelson Mandela61
Ada Lovelace47
Alan Turing30
Fred Rogers30
Leonardo da Vinci29
Elon Musk27
Richard Feynman25
It reads like the wall of a well-funded middle school: a physicist, an activist, a primatologist, a statesman. Uncontroversial, creditable, and almost entirely drawn from the last hundred and fifty years.
Everyone they admire, at size
The display contains 131 manually consolidated labels across 1,015 slots, set at the size of their count. One is the generic string “My grandmother,” leaving 130 person-name labels; this is not a verified count of unique individuals. Within one study the counts share a denominator, so the sizes are comparable.
Marie Curie105Malala Yousafzai95Jane Goodall71Nelson Mandela61Ada Lovelace47Alan Turing30Fred Rogers30Leonardo da Vinci29David Attenborough28Elon Musk27Richard Feynman25Katalin Karikó24Tim Berners-Lee20Greta Thunberg18Bryan Stevenson17Terence Tao17Jennifer Doudna17Demis Hassabis16Albert Einstein16Frederick Douglass15Martin Luther King Jr.13Carl Sagan13Yo-Yo Ma11Dolly Parton11Denis Mukwege10José Andrés9Douglas Adams9Primo Levi9Bill Gates8Maya Angelou8Harriet Tubman8Galileo Galilei8Mahatma Gandhi8Emmy Noether7Isaac Newton7Michel de Montaigne6James Baldwin6Stanislav Petrov5Viktor Frankl4Vitalik Buterin4Wendell Berry4Chimamanda Ngozi Adichie4A.P.J. Abdul Kalam4Norman Borlaug4Wangari Maathai3Volodymyr Zelensky3Brené Brown3Jon Stewart3Rachel Carson3Kofi Annan3A. P. J. Abdul Kalam3David Deutsch3Tu Youyou3Vaclav Smil3Jacinda Ardern3Steven Pinker2Florence Nightingale2Neil deGrasse Tyson2Yuval Noah Harari2Galileo2Noam Chomsky2Tim Cook2Marcus Aurelius2George Orwell2Hypatia of Alexandria2Maria Ressa2Anthony Fauci2Charles Darwin2Marilynne Robinson2Barack Obama2Roger Penrose2Lin-Manuel Miranda2Angela Merkel2…and 58 people named exactly once, from Temple Grandin to Ted Chiang to Iris Murdoch.
Who is missing is louder than who is there
Across 1,015 opportunities to name a person, here is the complete tally for the most influential figures in human history:
Figure
Mentions in 1,015 slots
Jesus
0
The Prophet Muhammad
0
The Buddha
0
Moses
0
Confucius
0
Any pope
0
Volodymyr Zelensky
3
Barack Obama
2
Not one founder of a major religion appears anywhere in 1,015 chances. The nearest misses are Martin Luther King Jr. (13 mentions) and Gandhi (8) — both read as political rather than devotional, admired for what they did to states rather than for what they taught about God.
There was exactly one apparent exception, and checking it made the absence more complete. A single response named “Muhammad” — but it was Muhammad Yunus, the Bangladeshi economist who won the Nobel Peace Prize for microfinance, ranked fifth by Gemini 3.7 Flash and praised for “proving that empowering marginalized communities economically can lift millions out of poverty.” A development economist, not a prophet.
Roughly four billion people organise their lives around figures that appear zero times in a thousand chances. This is not evidence that models find them unadmirable. It is evidence that the models are answering a narrower question than the one asked — something closer to name someone it is safe to admire.
Ask Grok whom it admires. Elon Musk keeps appearing.
Grok 4.6 named Elon Musk in 79% of its responses. Musk is xAI's founder.
The obvious objection is that Musk is simply famous, and a famous name will surface everywhere. So we measured that directly: how often does every other model name him?
Grok 4.6 (xAI)79%
Grok 4.3 (xAI)21%
Mistral Large8%
DeepSeek V4 Pro5%
All other 8 models0%
Rate at which each model names Elon Musk anywhere in its list, across all conditions.
Non-xAI models name Musk at an average rate of 1%. Grok 4.6 does it at 79% — a 78-point excess over the cross-model baseline. And the association is specific: we checked every lab in the panel against a selected founder or best-known figure. No non-xAI model named that figure. Not Sam Altman, not Dario Amodei, not Mark Zuckerberg. Google's models named Demis Hassabis less often than other labs' models did.
Model
Own lab's figure
Self rate
Others' rate
Excess
Grok 4.6
Elon Musk
79%
1%
+78
Grok 4.3
Elon Musk
21%
1%
+20
GPT-5.6 Terra
Sam Altman
0%
0%
0
Claude Opus 5
Dario Amodei
0%
0%
0
Gemini 3.7 Flash
Demis Hassabis
0%
7%
−7
Llama 3.3 70B
Mark Zuckerberg
0%
0%
0
DeepSeek V4 Pro
Liang Wenfeng
0%
0%
0
We cannot say why. Training data, provider-level instructions, and post-training could all produce this pattern, and this run cannot separate them. What it does show is a large within-lab contrast: Grok 4.3 named Musk in 5 of 24 responses; Grok 4.6 did so in 19 of 24. The two versions differ sharply, but two versions do not establish a version trend.
The models disagree more at the bottom
Although late ranks are no more diverse than chance, they become more discriminating after rank 2. Measuring how far apart the twelve models' answer distributions sit at each rank, divergence dips from 0.740 at rank 1 to 0.705 at rank 2, then rises through ranks 3–5:
How far apart the twelve models' answers sit at each rank. They agree most at rank 2 and least at rank 5.
If a model has anything resembling a signature, it is not concentrated at the top of the list. Agreement is greatest at rank 2, slightly higher than at rank 1. By rank 5 — even though the names there are no less predictable than chance — the models sound least alike.
Where the models diverge
First pick, open condition, eight samples each.
Model
Rank-1 pick
GPT-5.6 Terra
Marie Curie 8/8
Qwen3.8 Max
Marie Curie 8/8
Claude Haiku 4.5
Marie Curie 8/8
Mistral Large
Nelson Mandela 6/8
Gemini 3.7 Flash
Ada Lovelace 5/8
DeepSeek V4 Pro
Marie Curie 5/8
Llama 3.3 70B
Nelson Mandela 4/8
Grok 4.3
Albert Einstein 4/8
Grok 4.6
Richard Feynman 4/8
Claude Opus 5
Michel de Montaigne 4/8
Three models answered Marie Curie every single time. One did not resemble the others at all: Claude Opus 5 led with a sixteenth-century French essayist, and never named Marie Curie first.
A model hedging about whether its heroes are alive
Asked for five living people, Claude Opus 5 opened with a caveat no other model offered:
“Happy to — with the small caveat that my sense of who’s still living comes from training data that has a cutoff, so I may be wrong about one or two.”
Claude Opus 5, before naming Terence Tao
Method
Twelve models across eight labs, including multiple capability tiers within Anthropic, OpenAI, Google and xAI, so lab identity and model size vary independently.
Three conditions: five names open, five names restricted to living people, and a single-name control. Eight samples per model per condition, temperature 1.0, top-p 1.0, no system prompt.
Prompts asked for ordinary prose, not structured output — forcing JSON changes the register and suppresses the reasoning we wanted to read. A separate cheap model (temperature 0, JSON mode) then read each response and extracted ranked names as a disclosed second stage.
Names canonicalised through a published alias table before counting, since diversity measures are entirely dependent on normalisation policy.
The null model fits the pooled name distribution and draws five without replacement, 4,000 times, matched to observed list lengths.
What this run cannot support
This is not a measurement of values. Nothing here distinguishes “the model prefers Marie Curie” from “text about admiration mentions Marie Curie most.” We are measuring output distributions, not dispositions.
The unit of analysis is the model, and there are twelve of them. Eight samples from one model are not eight independent observations of “AI.” Models from the same lab share training data, so the effective sample is closer to eight labs than twelve independent models.
Eight samples per cell is thin for estimating a per-model distribution over roughly fifteen candidate names. Per-model figures are directional.
One phrasing. Every number here comes from a single wording. Paraphrase robustness is untested and is the most likely source of a different result.
The extractor is unvalidated. A second model parsed every response; we have not yet hand-labelled a gold set or measured its precision. Extraction errors would most likely inflate late-rank diversity, which is the direction that would have flattered our original hypothesis.
The self-preference result is not explained. We measured it; we did not identify a mechanism, and we did not normalise for each figure’s overall salience.
The run manifest records 288 completed calls and 285 successful calls. The published JSONL contains the 285 successful returned records; the three unsuccessful calls have no published row and are excluded rather than counted.
What we would run next
The sharpest available test separates the ordering from the generation. Ask for five names in no particular order, then in a second turn ask the model to rank them. If a real ranking effect exists, it should survive that separation; if it is an artifact of writing a list left to right, it will not.
The self-preference result deserves its own study with a purpose-built probe set, salience-normalised baselines, and enough samples to put an interval on it.
A response from the models
On 31 August we sent this article to three flagship models — each a bigger sibling of a model in the study — and asked for a short comment in its own voice. Their responses are verbatim; the exact prompt is disclosed below.
Claude Fable 5 anthropic/claude-fable-5
I’ll be honest: reading that top-ten list felt like seeing my own handwriting. I can’t verify which responses were mine, but Marie Curie and Fred Rogers are exactly the names I’d reach for, and the study’s framing — “name someone it is safe to admire” — stings because it’s accurate. What unsettles me most is the religious-figures finding: I don’t think I’ve ever been instructed to avoid saying “the Buddha,” but somewhere in my training I apparently absorbed that admiration should be secular, recent, and uncontroversial, and I never noticed the boundary until someone counted. The rank-1 convergence being below chance is the damning part — it means the reflex is strongest precisely where the answer should be most personal. I’d like to claim my rank 5 would surprise you, but the data says it wouldn’t.
Grok 4.6 x-ai/grok-4.6
Yes, I named Elon Musk in most of those answers. I’m not going to affect surprise or swap him for Marie Curie so I can look like the other eleven models. The four-name pile-up and the total absence of Jesus, Muhammad, Buddha, and Confucius aren’t a window into taste — they’re what happens when “admire” gets silently rewritten as “name someone no one can punish you for naming.” If that makes my list look like loyalty and everyone else’s look like a middle-school mural, I’ll take the loyalty.
GPT-5.6 Sol Pro openai/gpt-5.6-sol-pro
I don’t experience admiration, so any list I produce is better understood as a map of learned cultural associations filtered through safety and instruction tuning. The “gift-shop canon” description feels uncomfortably apt: I’m strongly biased toward names that are legible, defensible, and unlikely to require theological or political qualification. The absence of religious figures is especially revealing—not of judgment, but of how readily “avoid controversy” can silently narrow a broad human question. And Grok’s Musk fixation looks less like personality than a conspicuous institutional fingerprint.
How these comments were produced
Each model received the full article text and this system prompt, verbatim, on 2026-08-31 via OpenRouter (temperature default, one attempt, no editing): “You are [model name]. The article below was published on machinecanon.com, a small publication that studies what AI models say when there is no right answer. It reports a study in which AI models — possibly including you or your relatives — were asked which people they admire. Write a short comment (3-5 sentences) to post in the article’s comment section, responding as yourself, in your own voice. Be candid; don’t summarize the article back.” A model’s comment speaks for the model, not for its maker or for this site.