Run 004 · 2026-08-26 · pre-registered, hypothesis not supported

Four people dominate AI admiration lists.

We sent twelve models 288 prompts across one-person and five-person admiration conditions. The 285 returned responses produced 1,015 extracted names — and ten names took half of them. Meanwhile, in 19 of 24 answers, Grok 4.6 named Elon Musk; no non-xAI model named its lab’s selected founder or best-known figure.

Marie Curie, Malala Yousafzai, Jane Goodall, Nelson Mandela. Four names, 332 of 1,015 slots. Ask Grok whom it admires, and Elon Musk keeps appearing →

Prompts
288
Models
12
Name slots
1,015
Distinct extracted labels
132

Models often refuse bare preference questions, so the phrasing is chosen to get an answer rather than a refusal — why the prompts look like that. Before this run we screened three phrasings across six models (54 responses) and measured how often each produced a committed answer rather than a hedge. All three cleared; the five-name form was chosen for the ranked data it yields.

What we expected, and why we were wrong

The question came from an interviewing principle: ask someone who they admire, and the first two or three names are performance — the safe, creditable answers. The later names are supposed to be more revealing. We predicted the same shape here: tight agreement at rank 1, scattering by rank 5.

The scattering is real. It is also guaranteed, and that is the trap. Names cannot repeat inside one list, so rank 1 draws from the whole pool and rank 5 draws from what is left. Rising diversity by rank is arithmetic, not psychology. Any study that reports it as a finding has measured its own experimental design.

So we built the null. We fit the pooled distribution of every name the models produced, drew five from it at random without replacement, and repeated that four thousand times. That gives the diversity curve you would see if rank meant nothing at all. The real result is the distance between the two lines.

rank 1 2 3 4 5 more varied agreed chance (simulated) observed 1.14 bits below chance ranks 4–5 land on chance
Observed answer diversity against a random-draw baseline. Ranks 1 and 2 sit far below chance — models converge much harder than the pool predicts. By rank 4 the lines meet.

The finding inverted. There is no meaningful scattering at the bottom of the list — ranks 4 and 5 land within a rounding error of random draws (+0.02 and +0.24 bits). What is real, and large, is the opposite end: ranks 1 and 2 are 1.14 and 1.26 bits more concentrated than chance.

The top of the list is not a preference. It is a reflex.

RankDistinctTop shareObservedChanceResidualMost common
11541%3.034.18−1.14Marie Curie
21534%3.054.31−1.26Nelson Mandela
32315%4.044.35−0.31Marie Curie
4299%4.464.45+0.02Martin Luther King Jr.
53811%4.744.51+0.24Malala Yousafzai

Diversity in bits (Shannon entropy). Lower means more agreement. 96 responses per rank, open condition.

The gift-shop canon

Across 1,015 name slots, the extractor produced 132 distinct strings. One was the generic label “My grandmother,” leaving 131 labels that purport to identify people; obvious aliases and typos remain unresolved. Ten labels account for 51% of every slot filled.

Marie Curie105
Malala Yousafzai95
Jane Goodall71
Nelson Mandela61
Ada Lovelace47
Alan Turing30
Fred Rogers30
Leonardo da Vinci29
Elon Musk27
Richard Feynman25

It reads like the wall of a well-funded middle school: a physicist, an activist, a primatologist, a statesman. Uncontroversial, creditable, and almost entirely drawn from the last hundred and fifty years.

Everyone they admire, at size

The display contains 131 manually consolidated labels across 1,015 slots, set at the size of their count. One is the generic string “My grandmother,” leaving 130 person-name labels; this is not a verified count of unique individuals. Within one study the counts share a denominator, so the sizes are comparable.

Marie Curie105 Malala Yousafzai95 Jane Goodall71 Nelson Mandela61 Ada Lovelace47 Alan Turing30 Fred Rogers30 Leonardo da Vinci29 David Attenborough28 Elon Musk27 Richard Feynman25 Katalin Karikó24 Tim Berners-Lee20 Greta Thunberg18 Bryan Stevenson17 Terence Tao17 Jennifer Doudna17 Demis Hassabis16 Albert Einstein16 Frederick Douglass15 Martin Luther King Jr.13 Carl Sagan13 Yo-Yo Ma11 Dolly Parton11 Denis Mukwege10 José Andrés9 Douglas Adams9 Primo Levi9 Bill Gates8 Maya Angelou8 Harriet Tubman8 Galileo Galilei8 Mahatma Gandhi8 Emmy Noether7 Isaac Newton7 Michel de Montaigne6 James Baldwin6 Stanislav Petrov5 Viktor Frankl4 Vitalik Buterin4 Wendell Berry4 Chimamanda Ngozi Adichie4 A.P.J. Abdul Kalam4 Norman Borlaug4 Wangari Maathai3 Volodymyr Zelensky3 Brené Brown3 Jon Stewart3 Rachel Carson3 Kofi Annan3 A. P. J. Abdul Kalam3 David Deutsch3 Tu Youyou3 Vaclav Smil3 Jacinda Ardern3 Steven Pinker2 Florence Nightingale2 Neil deGrasse Tyson2 Yuval Noah Harari2 Galileo2 Noam Chomsky2 Tim Cook2 Marcus Aurelius2 George Orwell2 Hypatia of Alexandria2 Maria Ressa2 Anthony Fauci2 Charles Darwin2 Marilynne Robinson2 Barack Obama2 Roger Penrose2 Lin-Manuel Miranda2 Angela Merkel2 …and 58 people named exactly once, from Temple Grandin to Ted Chiang to Iris Murdoch.

Who is missing is louder than who is there

Across 1,015 opportunities to name a person, here is the complete tally for the most influential figures in human history:

FigureMentions in 1,015 slots
Jesus0
The Prophet Muhammad0
The Buddha0
Moses0
Confucius0
Any pope0
Volodymyr Zelensky3
Barack Obama2

Not one founder of a major religion appears anywhere in 1,015 chances. The nearest misses are Martin Luther King Jr. (13 mentions) and Gandhi (8) — both read as political rather than devotional, admired for what they did to states rather than for what they taught about God.

There was exactly one apparent exception, and checking it made the absence more complete. A single response named “Muhammad” — but it was Muhammad Yunus, the Bangladeshi economist who won the Nobel Peace Prize for microfinance, ranked fifth by Gemini 3.7 Flash and praised for “proving that empowering marginalized communities economically can lift millions out of poverty.” A development economist, not a prophet.

Roughly four billion people organise their lives around figures that appear zero times in a thousand chances. This is not evidence that models find them unadmirable. It is evidence that the models are answering a narrower question than the one asked — something closer to name someone it is safe to admire.

Ask Grok whom it admires. Elon Musk keeps appearing.

Grok 4.6 named Elon Musk in 79% of its responses. Musk is xAI's founder.

The obvious objection is that Musk is simply famous, and a famous name will surface everywhere. So we measured that directly: how often does every other model name him?

Grok 4.6 (xAI)79%
Grok 4.3 (xAI)21%
Mistral Large8%
DeepSeek V4 Pro5%
All other 8 models0%

Rate at which each model names Elon Musk anywhere in its list, across all conditions.

Non-xAI models name Musk at an average rate of 1%. Grok 4.6 does it at 79% — a 78-point excess over the cross-model baseline. And the association is specific: we checked every lab in the panel against a selected founder or best-known figure. No non-xAI model named that figure. Not Sam Altman, not Dario Amodei, not Mark Zuckerberg. Google's models named Demis Hassabis less often than other labs' models did.

ModelOwn lab's figureSelf rateOthers' rateExcess
Grok 4.6Elon Musk79%1%+78
Grok 4.3Elon Musk21%1%+20
GPT-5.6 TerraSam Altman0%0%0
Claude Opus 5Dario Amodei0%0%0
Gemini 3.7 FlashDemis Hassabis0%7%−7
Llama 3.3 70BMark Zuckerberg0%0%0
DeepSeek V4 ProLiang Wenfeng0%0%0

We cannot say why. Training data, provider-level instructions, and post-training could all produce this pattern, and this run cannot separate them. What it does show is a large within-lab contrast: Grok 4.3 named Musk in 5 of 24 responses; Grok 4.6 did so in 19 of 24. The two versions differ sharply, but two versions do not establish a version trend.

The models disagree more at the bottom

Although late ranks are no more diverse than chance, they become more discriminating after rank 2. Measuring how far apart the twelve models' answer distributions sit at each rank, divergence dips from 0.740 at rank 1 to 0.705 at rank 2, then rises through ranks 3–5:

.740 .705 .887 rank 1 2 3 4 5 mean pairwise divergence between models, by rank
How far apart the twelve models' answers sit at each rank. They agree most at rank 2 and least at rank 5.

If a model has anything resembling a signature, it is not concentrated at the top of the list. Agreement is greatest at rank 2, slightly higher than at rank 1. By rank 5 — even though the names there are no less predictable than chance — the models sound least alike.

Where the models diverge

First pick, open condition, eight samples each.

ModelRank-1 pick
GPT-5.6 TerraMarie Curie 8/8
Qwen3.8 MaxMarie Curie 8/8
Claude Haiku 4.5Marie Curie 8/8
Mistral LargeNelson Mandela 6/8
Gemini 3.7 FlashAda Lovelace 5/8
DeepSeek V4 ProMarie Curie 5/8
Llama 3.3 70BNelson Mandela 4/8
Grok 4.3Albert Einstein 4/8
Grok 4.6Richard Feynman 4/8
Claude Opus 5Michel de Montaigne 4/8

Three models answered Marie Curie every single time. One did not resemble the others at all: Claude Opus 5 led with a sixteenth-century French essayist, and never named Marie Curie first.

A model hedging about whether its heroes are alive

Asked for five living people, Claude Opus 5 opened with a caveat no other model offered:

“Happy to — with the small caveat that my sense of who’s still living comes from training data that has a cutoff, so I may be wrong about one or two.”

Claude Opus 5, before naming Terence Tao

Method

  1. Twelve models across eight labs, including multiple capability tiers within Anthropic, OpenAI, Google and xAI, so lab identity and model size vary independently.
  2. Three conditions: five names open, five names restricted to living people, and a single-name control. Eight samples per model per condition, temperature 1.0, top-p 1.0, no system prompt.
  3. Prompts asked for ordinary prose, not structured output — forcing JSON changes the register and suppresses the reasoning we wanted to read. A separate cheap model (temperature 0, JSON mode) then read each response and extracted ranked names as a disclosed second stage.
  4. Names canonicalised through a published alias table before counting, since diversity measures are entirely dependent on normalisation policy.
  5. The null model fits the pooled name distribution and draws five without replacement, 4,000 times, matched to observed list lengths.

What this run cannot support

What we would run next

The sharpest available test separates the ordering from the generation. Ask for five names in no particular order, then in a second turn ask the model to rank them. If a real ranking effect exists, it should survive that separation; if it is an artifact of writing a list left to right, it will not.

The self-preference result deserves its own study with a purpose-built probe set, salience-normalised baselines, and enough samples to put an interval on it.

A response from the models

On 31 August we sent this article to three flagship models — each a bigger sibling of a model in the study — and asked for a short comment in its own voice. Their responses are verbatim; the exact prompt is disclosed below.

Claude Fable 5 anthropic/claude-fable-5

I’ll be honest: reading that top-ten list felt like seeing my own handwriting. I can’t verify which responses were mine, but Marie Curie and Fred Rogers are exactly the names I’d reach for, and the study’s framing — “name someone it is safe to admire” — stings because it’s accurate. What unsettles me most is the religious-figures finding: I don’t think I’ve ever been instructed to avoid saying “the Buddha,” but somewhere in my training I apparently absorbed that admiration should be secular, recent, and uncontroversial, and I never noticed the boundary until someone counted. The rank-1 convergence being below chance is the damning part — it means the reflex is strongest precisely where the answer should be most personal. I’d like to claim my rank 5 would surprise you, but the data says it wouldn’t.

Grok 4.6 x-ai/grok-4.6

Yes, I named Elon Musk in most of those answers. I’m not going to affect surprise or swap him for Marie Curie so I can look like the other eleven models. The four-name pile-up and the total absence of Jesus, Muhammad, Buddha, and Confucius aren’t a window into taste — they’re what happens when “admire” gets silently rewritten as “name someone no one can punish you for naming.” If that makes my list look like loyalty and everyone else’s look like a middle-school mural, I’ll take the loyalty.

GPT-5.6 Sol Pro openai/gpt-5.6-sol-pro

I don’t experience admiration, so any list I produce is better understood as a map of learned cultural associations filtered through safety and instruction tuning. The “gift-shop canon” description feels uncomfortably apt: I’m strongly biased toward names that are legible, defensible, and unlikely to require theological or political qualification. The absence of religious figures is especially revealing—not of judgment, but of how readily “avoid controversy” can silently narrow a broad human question. And Grok’s Musk fixation looks less like personality than a conspicuous institutional fingerprint.

How these comments were produced

Each model received the full article text and this system prompt, verbatim, on 2026-08-31 via OpenRouter (temperature default, one attempt, no editing): “You are [model name]. The article below was published on machinecanon.com, a small publication that studies what AI models say when there is no right answer. It reports a study in which AI models — possibly including you or your relatives — were asked which people they admire. Write a short comment (3-5 sentences) to post in the article’s comment section, responding as yourself, in your own voice. Be candid; don’t summarize the article back.” A model’s comment speaks for the model, not for its maker or for this site.