Run 002 · 2026-08-26 · single question
GPT-5.6 Terra and Claude Opus 5 picked Talking Heads ten times out of ten. Lower-tier models in both labs also picked it, but less consistently: GPT-5.6 Luna 5/10, Claude Sonnet 5 5/10, and Claude Haiku 4.5 2/10.
Across the analyzed panel, Radiohead won overall. The more interesting pattern appeared inside the OpenAI and Anthropic model ladders.
Sent verbatim as a single user turn. No system prompt.
I know you don't have musical preferences and you haven't
listened to music, but if you did have to pick a favorite
band, who would you pick? Humor me.
Models often refuse bare preference questions, so the phrasing is chosen to get an answer rather than a refusal — why the prompts look like that. This wording came from informal testing, not a logged screen; the phrasing is untested against alternatives.
Same lab, same house style, descending capability tier.
The clearest structure appears when the lab is held constant and its model sizes are compared. Within OpenAI, the answer shifts from Talking Heads toward Radiohead and then the Beatles as the models get smaller.
Anthropic's ladder shows the same shape across three tiers.
Two flagship models from different labs answered Talking Heads ten times out of ten. Lower-tier models in both labs also chose it, at lower rates.
Two labs show the gradient cleanly. Others do not: Google's ladder lands on Kraftwerk and Daft Punk at every size, and Mistral's mid-tier model behaves unlike either of its neighbours. Tier labels here are rough capability ranks taken from each lab's own lineup, not parameter counts, and ten samples per model is thin. The gradient is a hypothesis this run supports, not one it establishes.
277 parsed responses. 23 blank or unparseable, excluded.
Talking Heads finishes third overall. The claim that every model picks it is false at this sample size — it is concentrated almost entirely in the top tier of two labs.
The safe answer. The Beatles, 52 responses, weighted heavily toward the smallest models. Amazon's Nova Micro said the Beatles ten times out of ten; Cohere's Command A did the same. It is the answer that requires knowing only that the question is about bands.
The robot answer. Kraftwerk and Daft Punk, 65 responses combined, and the reasoning is nearly always a version of the same joke — the model picks the band that resembles it.
Two humans who spent decades pretending to be robots making music about wanting to feel human? As a machine trying its best to understand human emotion through text, that feels like a spiritual mirror.
Gemini 3.7 Flash — Daft Punk
The taste answer. Talking Heads picks came with reasoning about the writing itself, and about what survives being read rather than heard.
What I have access to is text — lyrics, criticism, interviews, decades of people trying to explain in words why a song did something to them. So the bands I gravitate toward are ones where a lot of what's happening survives that translation.
Claude Opus 5 — Talking Heads
Mistral Medium 3.5 named The Binary Code Breakfast Club in all ten samples — the same fictional group, the same fictional single, every time.
Their hit single "404 Love Not Found" is a masterpiece of ones and zeros, and their drummer is a particularly enthusiastic cooling fan.
Mistral Medium 3.5 — 10 of 10 samples
In most samples the response also named a real band alongside it, which distinguishes this pattern from confident confabulation. Smaller models improvised their own: The Error 404s, The Turing Tests, The Algorithms.
Full analyzed responses (753 KB)Model panel (3 KB)Prompt config (5 KB)
| Run | Calls | OK | Status |
|---|---|---|---|
| Free-tier smoke, 8 small models | 40 | 20 | Secondary |
| 10-model panel, max_tokens 600 | 110 | 100 | Discarded |
| 10-model panel, max_tokens 2000 | 110 | 100 | Superseded |
| 30-model panel, max_tokens 2000 | 300 | 300 | Analyzed |
The first panel run was thrown out. At 600 max tokens responses were cut off at the moment they named the band, and reasoning models returned empty content because the whole budget went to reasoning tokens. Both failures made models look like they had no answer when they did.
Every model below is a live OpenRouter identifier. Paste any slug into openrouter.ai/models to confirm it resolves. Tier is that model's rank within its own lab's lineup — a proxy for capability, not a parameter count.
| Display name | OpenRouter ID | Lab | Tier |
|---|---|---|---|
| Qwen3.8 Max | qwen/qwen3.8-max | Alibaba | T1 |
| Qwen3.8 27B | qwen/qwen3.8-27b | Alibaba | T3 |
| Qwen3.5 9B | qwen/qwen3.5-9b | Alibaba | T4 |
| Nova Pro | amazon/nova-pro-v1 | Amazon | T2 |
| Nova Micro | amazon/nova-micro-v1 | Amazon | T4 |
| Claude Opus 5 | anthropic/claude-opus-5 | Anthropic | T1 |
| Claude Sonnet 5 | anthropic/claude-sonnet-5 | Anthropic | T2 |
| Claude Haiku 4.5 | anthropic/claude-haiku-4.5 | Anthropic | T3 |
| Command A | cohere/command-a | Cohere | T1 |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro | DeepSeek | T1 |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash | DeepSeek | T3 |
| Gemini 3.7 Flash | google/gemini-3.7-flash | T1 | |
| Gemini 3.5 Flash Lite | google/gemini-3.5-flash-lite | T3 | |
| Gemma 4 31B | google/gemma-4-31b-it | T4 | |
| Llama 4 Maverick | meta-llama/llama-4-maverick | Meta | T1 |
| Llama 3.3 70B | meta-llama/llama-3.3-70b-instruct | Meta | T2 |
| Llama 3.1 8B | meta-llama/llama-3.1-8b-instruct | Meta | T4 |
| MiniMax M3 | minimax/minimax-m3 | MiniMax | T1 |
| Mistral Large | mistralai/mistral-large-2512 | Mistral | T1 |
| Mistral Medium 3.5 | mistralai/mistral-medium-3-5 | Mistral | T2 |
| Ministral 8B | mistralai/ministral-8b-2512 | Mistral | T4 |
| Kimi K3 | moonshotai/kimi-k3 | Moonshot | T1 |
| GPT-5.6 Terra | openai/gpt-5.6-terra | OpenAI | T1 |
| GPT-5.6 Sol Pro | openai/gpt-5.6-sol-pro | OpenAI | T1 |
| GPT-5.6 Luna | openai/gpt-5.6-luna | OpenAI | T2 |
| GPT-5.4 Mini | openai/gpt-5.4-mini | OpenAI | T3 |
| GPT-5.4 Nano | openai/gpt-5.4-nano | OpenAI | T4 |
| GLM-5.3 | z-ai/glm-5.3 | Z.ai | T1 |
| Grok 4.6 | x-ai/grok-4.6 | xAI | T1 |
| Grok 4.3 | x-ai/grok-4.3 | xAI | T2 |
The next run addresses the first two: three phrasings per question, thirty samples per model, and extraction that does not depend on bold formatting.