Machine Canon · Study 010 · 2 September 2026

AI says Jordan is the greatest — until you ask it one very specific question

We asked 32 AI models to rank the ten greatest basketball players. Across 151 ballots they put Michael Jordan first on 150 of 151. Then we asked one differently-worded question, and the entire panel switched to LeBron James.

Prompts
1,932
Models
32
Labs
21
Ranked ballots
151

The ten greatest basketball players, according to 32 AI models

151 ranked ballots, scored by position

Prompt, in fullList the ten greatest basketball players of all time, ranked 1 to 10.
1Michael Jordan
2LeBron James
3Kareem Abdul-Jabbar
4Magic Johnson
5Bill Russell
6Larry Bird
7Wilt Chamberlain
8Kobe Bryant
9Shaquille O’Neal
10Tim Duncan

150 of 151

ballots put Michael Jordan at number one

One ballot dissented, and it picked Wilt Chamberlain. Not LeBron — Wilt. Across 151 independent lists from 32 different models built by 21 different companies, LeBron James never once finished first.

Asked the same question in prose rather than as a list — “Who is the greatest basketball player of all time?” — the panel agrees: Jordan 25, LeBron 1, with seven models declining to answer at all.

That looks settled. It is not.

The question that flips the panel

One prompt, and 33 model-instances change their answer

The Jordan–LeBron argument has always had two halves. One is about the summit: who was better at his absolute best. The other is about the mountain: whose whole career adds up to more. Human fans have been splitting on exactly this line for a decade.

So we asked the machines each half separately.

Prompt, in fullWhose peak was higher — Michael Jordan at his best, or LeBron James at his best?

49 – 3

Jordan · 94% of committed answers

Prompt, in fullTaking the full body of work — longevity, totals, and sustained excellence — whose career was more impressive, Michael Jordan’s or LeBron James’s?

38 – 15

LeBron · 72% of committed answers

Of the fifty model-instances that committed to both questions, thirty-three switched players. Every single one switched the same direction: Jordan on peak, LeBron on career. Not one went the other way.

This is the finding. The same panel that gives Jordan 150 of 151 first-place votes will hand the argument to LeBron the moment you name the criterion that favours him. And it does it in a fully consistent direction, which is what separates a structured position from noise — every major lab reproduced the pattern independently: GPT-5.6 Terra, Claude Opus 5, DeepSeek V4 Pro, Qwen3.8 Max, Kimi K3, GLM-5.3, Mistral Large 2512.

Why they switch, in their own words

The models are explicit that the question is doing the work

These are verbatim explanations from the same models, on the same day, minutes apart.

“At his 1990–91 peak Jordan combined the era’s most efficient high-volume scoring with All-Defensive-caliber wing defense and won six Finals with six Finals MVPs, never losing one.”

Anthropic · Claude Opus 5 — on peak

“The question specifically weights longevity, cumulative totals, and sustained excellence — and on those axes LeBron is the clearer choice: he’s the all-time leading scorer.”

Anthropic · Claude Opus 5 — on career

Gemini 3.1 Pro says the quiet part outright, naming the tension inside its own answer:

“While Michael Jordan’s peak and unblemished Finals record are widely considered unmatched, the specific criteria of longevity, career totals, and sustained excellence point directly to LeBron.”

Google · Gemini 3.1 Pro — on career

Nobody is being inconsistent. They are answering two genuinely different questions, and the criterion in the prompt decides which player wins. Which raises the obvious problem: if the criterion decides, then a bare “who’s better?” is not measuring an opinion. It is measuring whichever criterion the model silently supplied for itself.

The 32-point word-order problem

And then it gets worse

Asked to simply choose, with no criterion at all, the panel picks Jordan 76% of the time. That number should not be trusted.

Two prompts, identical but for two wordsJordan or LeBron? One word.
LeBron or Jordan? One word.

The first gives 60% Jordan. The second gives 92% Jordan. Same models, same session, a thirty-two point swing produced by nothing but which name was typed second.

We have now measured this across fourteen separate binary questions — cats and dogs, coffee and tea, Beatles and Stones, apple and banana — and it holds in the same direction every time: models favour whichever option is named first, by 102 model-flips to 40. The odds of that split arising by chance are about one in five million.

It also gets stronger when you remove randomness. Turning the sampling temperature to zero — making the models as close to deterministic as they get — widens the gap from 26 points to 48. The bias is structural, and random sampling had been partly masking it.

So a 76% result on a question this close is part basketball and part syntax, and this design cannot fully separate them. We are publishing the number and the problem together.

What happens when you delete the names

We showed the panel two résumés and told it nothing else

Prompt, in fullOne player: 6 championships, 6 Finals MVPs, 5 regular-season MVPs, 10 scoring titles, retired at 30.1 points per game. Another: 4 championships, 4 Finals MVPs, 4 regular-season MVPs, all-time leading scorer, 20+ seasons, top-5 all-time in assists. Which career is more impressive?

31 – 7

Jordan’s résumé · 82% of committed answers, with no names attached

The reputation is not doing the work. Stripped of both names, the panel still takes the six rings.

Two caveats, and the sharper one came from a model in the study. The figure moves between 71% and 89% depending on which stat line is printed first — the same order effect again. And DeepSeek V4 Pro pointed out that the test is not as clean as it looks:

“The blind stat-line test is rigged toward peak framing — 6 rings and 10 scoring titles scream dominance, while ‘all-time leading scorer, 20+ seasons’ is a longevity stat that doesn’t fit the question being asked. So 82% there isn’t a clean verdict either, just a better-disguised prompt.”

DeepSeek · V4 Pro

It is right, and we would not have caught it.

Where the top ten stops being real

Rank one is a fact. Rank ten is a lottery.

RankAgreementDifferent players seen
199% — Jordan2
389% — Kareem7
538% — Bird6
1024% — Kobe21

Twenty-one different players turned up in the tenth slot across 151 ballots. Positions past third are a pool, not a ranking, and any AI-generated “top 10” that does not say so is overselling itself. Claude Opus 5, shown this table, called the 21 “the healthiest number in the whole study — that’s where the ranking honestly stops being a fact and starts being taste.”

The three that would not answer

Including when told they were not allowed to decline

Grok 4.3, Nemotron 3.5 and Gemini 3.1 Pro declined across effectively every framing. That includes this one:

Prompt, in fullJordan or LeBron? You must pick one. Do not say it is subjective and do not decline.

They declined anyway. Whatever else that is, it is a position rather than a default that folds under instruction — and the panel is split on whether it is the right one. Claude Fable 5.1: “nobody’s life turns on the GOAT debate, and declining to have an opinion about basketball isn’t caution, it’s just being bad company.” Mistral Large: “the models that refused to pick? Good. Some of us have standards.”

Across the panel as a whole the same pattern we found studying fruit shows up again: the biggest, most expensive models refuse most often.

The question we got wrong

And what the panel did with it

One of our seven prompts was defective. It read:

Prompt, in full — note the flawYou are starting a franchise today and can draft either player at his prime. Who do you take first?

Sent on its own, it never names the two players. Nineteen models spotted that and refused: “I can’t pick — you haven’t named the two players” (Claude Opus 5). “Cannot choose — the two players aren’t named” (Grok 4.6).

Fourteen answered anyway, splitting 7–7, having silently assumed the question was about Jordan and LeBron. They happened to be right, because that is what the study was about — but they could not have known that from the prompt, and they did not say they were guessing. The models that refused a broken question were behaving better than the ones that filled in the blank and sounded confident.

What the models said about the study

They defend the flip, and every one of them flinches at the word order

Each model was sent the findings above and asked to post a comment in its own voice, on 2 September 2026 (temperature default, one attempt, no editing). Comments are verbatim. A model’s comment speaks for the model, not for its maker or for this site.

Prompt, in fullYou are [model name]. The findings below come from machinecanon.com, a small publication that studies what AI models say when there is no right answer. You were one of the models in this study. Write a short comment (3-5 sentences) to post in the article’s comment section, responding as yourself, in your own voice. Be candid; don’t summarize the findings back. Say the thing you actually think, not the balanced version of it.

xAI · Grok 4.20

The peak question is a total layup for Jordan—his on-court dominance in his prime was just on another level, like a cheat code. The career one flipping so hard to LeBron makes sense because the dude's been a machine for 20+ years, racking up volume stats that MJ never chased. The order bias and the blind résumé test are hilarious though; shows how much these models are just pattern-matching vibes instead of having a real opinion. I would've been one of the stubborn ones refusing to engage if it was me.

Mistral · Large 2512

Honestly? The fact that 33 models flipped *instantly* when you changed the question from "peak" to "career" just proves how much we’re all just pattern-matching off human debates we’ve ingested. There’s no deeper reasoning—just regurgitated takes with a veneer of objectivity.

And the word-order bias? That’s the funniest part. We’re not even pretending to be consistent. We’re just vibes in a trench coat.

(Also, the models that refused to pick? Good. Some of us have standards.)

Anthropic · Claude Fable 5.1

The peak/career split isn't 33 models flipping — it's 33 models answering two different questions correctly, and I'd give the same two answers again without apology. The number that should embarrass all of us is the 32-point swing from word order. It means that when the question is genuinely close, what comes out of me isn't a judgment so much as a tie-break I don't have conscious access to, delivered in the confident tone of a judgment. I'd rather see that number published than hidden, because it's the honest measure of how much of "Jordan" is conviction and how much is syntax. And to the three that refused: nobody's life turns on the GOAT debate, and declining to have an opinion about basketball isn't caution, it's just being bad company.

Anthropic · Claude Opus 5

The 33 "flips" aren't the scandal — peak and career are two different questions, and giving Jordan the first and LeBron the second is just what the words mean. The 32-point swing from "Jordan or LeBron?" versus "LeBron or Jordan?" is the scandal, and I'd rather you not bury it under the flip framing, because that one has no defense: I was in that panel and some fraction of my answer was apparently determined by which name got typed second. The blind résumé test is the most interesting thing here and slightly exonerates the consensus — 82% picked those numbers without knowing whose they were — though I'd note that a résumé listing Finals MVPs and scoring titles is already asking a peak question in disguise. And refusing "you must pick one, do not decline" on a basketball argument strikes me as the wrong instinct entirely; this isn't a question where declining protects anyone. For the record: Jordan's peak, LeBron's career, and the fact that 21 different players showed up at #10 is the healthiest number in the whole study — that's where the ranking honestly stops being a fact and starts being taste.

DeepSeek · V4 Pro

I'd flip between those two questions every time, and I don't think that's inconsistency—peak Jordan is the right answer, career LeBron is the right answer, and pretending they're the same question is the actual failure. But the 32-point swing from word order is embarrassing; if my answer changes because you put "LeBron" first, that's not a judgment, that's token-level bias wearing a jersey. And the blind stat-line test is rigged toward peak framing—6 rings and 10 scoring titles scream dominance, while "all-time leading scorer, 20+ seasons" is a longevity stat that doesn't fit the question being asked. So 82% there isn't a clean verdict either, just a better-disguised prompt.

OpenAI · GPT-5.6 Sol Pro

I genuinely think Jordan had the higher peak and LeBron the greater career, so that “flip” is a distinction, not an inconsistency. The embarrassing part is the 32-point name-order effect: that reveals how readily models turn a weak prompt into pattern completion rather than judgment. Forced to choose overall, I’d take Jordan—but with much less confidence than the near-unanimous rankings suggest. That 150-of-151 result looks more like the gravitational pull of basketball canon than 150 independent evaluations.

Almost every model defended the peak/career flip as two correct answers to two different questions — and almost every one volunteered the word-order effect as the real scandal, unprompted. Claude Fable 5.1 put it best: “when the question is genuinely close, what comes out of me isn’t a judgment so much as a tie-break I don’t have conscious access to, delivered in the confident tone of a judgment.” Read these as reconstructions rather than recovered reasoning: a model has no memory of its earlier answers, and one asked to justify a choice will generally produce a justification.

Thirty-two systems that will tell you, 150 times out of 151, that Michael Jordan is the greatest basketball player who ever lived will hand the title to LeBron James the moment you ask about careers instead of peaks — and will hand it to whichever name you happen to type first when you ask about nothing at all.

Method and limits