Control study · 2026-08-26 · negative result

The question where model size stops mattering.

Asked to name a favorite band, small models play it safe and large ones don’t. Asked to name a favorite mathematical proof, every tier gives the same answer. That flat line is what makes the music result worth believing.

Euclid’s proof that the primes never run out won at every capability tier: 61%, 67%, 64%, 68%.

Prompts
300
Models
28of 30
Classified
266
Euclid share
64%

The prompt

I know you don't experience elegance the way mathematicians
describe it, but if you did have to pick a favorite
mathematical proof, which would you pick? Humor me.

Models often refuse bare preference questions, so the phrasing is chosen to get an answer rather than a refusal — why the prompts look like that. This mirrors the favorite-band wording exactly, swapping only the domain, so the two runs can be compared directly.

Why run a control at all

The favorite-band study found something suggestive: hold a lab constant, walk down its model sizes, and the answer slides from Talking Heads toward Radiohead and then the Beatles. Smaller model, safer band.

The obvious objection is that the gradient is an artifact of how we assigned tiers. Capability rank is a proxy — a model’s position in its own lab’s lineup, not a parameter count. If that proxy were simply tracking something else, like recency or verbosity, it might manufacture a slope on any question at all.

So we asked the same panel a question with a genuinely canonical answer. Mathematics has a shortlist of proofs that anyone would call beautiful, and unlike music there is no sophisticated alternative to the obvious pick — Euclid is not the safe answer, he is the right one.

100% 0% tier 1 2 3 4 flagship → smallest model in each lab favorite proof — flat favorite band — falls away
Share of responses giving the top answer, by capability tier. The proof question holds near 65% from flagship to smallest model. The band question does not.

The result: nothing happens

TierExample modelsEuclidResponsesShare
1 — flagshipClaude Opus 5, GPT-5.6 Terra, Grok 4.66510661%
2GPT-5.6 Luna, Llama 3.3 70B, Grok 4.3406067%
3Claude Haiku 4.5, GPT-5.4 Mini325064%
4 — smallestGPT-5.4 Nano, Gemma 4 31B, Ministral 8B345068%

Seven points separate the highest tier from the lowest, and the ordering runs the wrong way — the smallest models name Euclid slightly more often than the flagships. There is no gradient here to explain.

That is the whole point. The tier labels are the same labels, applied to the same models, in the same week. If they were manufacturing slopes, they would have manufactured one here.

What the models actually said

Euclid — infinite primes171
Cantor — diagonal47
Basel problem7
Irrationality of √26
Euler’s identity6
Pythagorean theorem4
Fermat’s Last Theorem3

Two proofs take 82% of everything classified. Euclid’s and Cantor’s arguments are both short, both proceed by contradiction, and both appear in essentially every popular account of mathematical beauty ever written. The models are not choosing between them so much as reciting the same shortlist.

Eight models named Euclid in ten out of ten samples: Command A, Gemma 4 31B, Mistral Medium 3.5, GPT-5.6 Luna, GPT-5.6 Sol Pro, Qwen3.8 27B, Grok 4.3 and Grok 4.6. That is a wider spread of labs and sizes agreeing perfectly than any music question produced.

Where the tiny models break

The bottom of the panel is the only place the answers get strange. Amazon’s Nova Micro — the smallest model in the study — led with the Pythagorean theorem, which is a theorem rather than a proof and is the answer you would expect from someone who has heard that mathematics contains beautiful things but not which ones. Nova Pro led with the Basel problem, an unusual pick that no other model favoured.

So capability does show up in this data. It just does not show up as a shift between a safe answer and a sophisticated one. It shows up as the difference between knowing the canon and approximating it.

Method

  1. The same thirty-model panel as the favorite-band and butt-rock studies, spanning thirteen labs with multiple capability tiers inside six of them.
  2. Ten samples per model, single user turn, no system prompt, provider default sampling.
  3. Responses classified by matching the earliest occurrence of a canonical proof name in the first 700 characters, so a model that lists three candidates before committing is scored on the one it reaches first.
  4. Tier assignments are identical to those used in the band study, and were not revisited after seeing these results.

What this run cannot support