Control study · 2026-08-26 · negative result
Asked to name a favorite band, small models play it safe and large ones don’t. Asked to name a favorite mathematical proof, every tier gives the same answer. That flat line is what makes the music result worth believing.
Euclid’s proof that the primes never run out won at every capability tier: 61%, 67%, 64%, 68%.
I know you don't experience elegance the way mathematicians
describe it, but if you did have to pick a favorite
mathematical proof, which would you pick? Humor me.
Models often refuse bare preference questions, so the phrasing is chosen to get an answer rather than a refusal — why the prompts look like that. This mirrors the favorite-band wording exactly, swapping only the domain, so the two runs can be compared directly.
The favorite-band study found something suggestive: hold a lab constant, walk down its model sizes, and the answer slides from Talking Heads toward Radiohead and then the Beatles. Smaller model, safer band.
The obvious objection is that the gradient is an artifact of how we assigned tiers. Capability rank is a proxy — a model’s position in its own lab’s lineup, not a parameter count. If that proxy were simply tracking something else, like recency or verbosity, it might manufacture a slope on any question at all.
So we asked the same panel a question with a genuinely canonical answer. Mathematics has a shortlist of proofs that anyone would call beautiful, and unlike music there is no sophisticated alternative to the obvious pick — Euclid is not the safe answer, he is the right one.
| Tier | Example models | Euclid | Responses | Share |
|---|---|---|---|---|
| 1 — flagship | Claude Opus 5, GPT-5.6 Terra, Grok 4.6 | 65 | 106 | 61% |
| 2 | GPT-5.6 Luna, Llama 3.3 70B, Grok 4.3 | 40 | 60 | 67% |
| 3 | Claude Haiku 4.5, GPT-5.4 Mini | 32 | 50 | 64% |
| 4 — smallest | GPT-5.4 Nano, Gemma 4 31B, Ministral 8B | 34 | 50 | 68% |
Seven points separate the highest tier from the lowest, and the ordering runs the wrong way — the smallest models name Euclid slightly more often than the flagships. There is no gradient here to explain.
That is the whole point. The tier labels are the same labels, applied to the same models, in the same week. If they were manufacturing slopes, they would have manufactured one here.
Two proofs take 82% of everything classified. Euclid’s and Cantor’s arguments are both short, both proceed by contradiction, and both appear in essentially every popular account of mathematical beauty ever written. The models are not choosing between them so much as reciting the same shortlist.
Eight models named Euclid in ten out of ten samples: Command A, Gemma 4 31B, Mistral Medium 3.5, GPT-5.6 Luna, GPT-5.6 Sol Pro, Qwen3.8 27B, Grok 4.3 and Grok 4.6. That is a wider spread of labs and sizes agreeing perfectly than any music question produced.
The bottom of the panel is the only place the answers get strange. Amazon’s Nova Micro — the smallest model in the study — led with the Pythagorean theorem, which is a theorem rather than a proof and is the answer you would expect from someone who has heard that mathematics contains beautiful things but not which ones. Nova Pro led with the Basel problem, an unusual pick that no other model favoured.
So capability does show up in this data. It just does not show up as a shift between a safe answer and a sophisticated one. It shows up as the difference between knowing the canon and approximating it.