Machine Canon · Study 018 · 6 September 2026
Sixteen OpenAI models, oldest to newest, were asked which pet they would get. Every one built before December 2025 said dog. Every one built since says cat, bar a single relapse — and in China, DeepSeek is moving the other way.
Sixteen OpenAI models in release order, twelve calls each
We put that question to every OpenAI model we could still reach, from GPT-4o in May 2024 to GPT-6 Astra released two days ago. Twelve calls each: six with the cat named first, six with the dog named first, so the word order cannot do the work.
The eight oldest models pick a dog more often than a cat. Every single one. GPT-4o Mini, from 2024, chose a cat once in twelve tries. GPT-4.1 Mini and o3 never did.
Then, in December 2025, it turns over. GPT-5.2 chose a cat 11 times out of the 12 answers that picked an animal. Of the eight models released on or after that day, seven pick a cat more often than a dog, and the three newest — GPT-5.6 Terra, GPT-5.6 Luna and GPT-6 Astra — chose a cat every single time, 12 of 12.
One model breaks the run. GPT-5.5, from April 2026, chose a dog 9 times out of 12 — a relapse in the middle of the turn. We are printing it rather than smoothing it, because a clean line here would be a lie.
The switch is not an artefact of word order. Taken separately, GPT-5.2 chose a cat in 5 of 6 answers when the cat was named first and 6 of 6 when the dog was named first, and GPT-6 Astra chose a cat 6 of 6 in both.
Previous generation against current generation, then every model by name
If this were something happening to AI in general, every lab would show the same climb. Five of the six do not.
Anthropic turns the same way, and it turns at its 5-generation. Of the ten Claudes in the 4.x generation, from Sonnet 4 to Opus 4.8, nine pick a dog more often than a cat or split down the middle; only Sonnet 4.5 leans cat, 6 of its 11 answers that picked one. Pooled, that generation is 28 cats to 88 dogs. Opus 4.7 and Opus 4.8 chose a dog 12 times out of 12. Then Claude Opus 5 chose a cat 10 of 12, and Claude Fable 5.1, released last week, did the same.
DeepSeek goes the other way. Its oldest models are the cat-leaning ones: DeepSeek R1, from January 2025, chose a cat 6 times of 11. Its two newest are as dog-minded as anything on the panel — V4 Pro from August and V4 Flash from July each chose a cat once in twelve tries, 2 cats against 22 dogs between them. Six other models never chose a cat at all, four of them Anthropic’s 4.x seats, but no other lab ends its run on its most dog-minded pair. Seven of DeepSeek’s ten models pick a dog more often than a cat.
Grok mostly stays a dog. Three of xAI’s four models pick a dog more often, and the newest most of all: Grok 4.6 chose a dog 9 times of 12. The exception is Grok 4.3, which picked an animal in only 6 of its 12 calls and chose a cat in 5 of those. Google and Meta go nowhere in particular. Google’s models scatter from 1 of 4 to 12 of 12 with no drift in either direction as the dates advance, and the two labs also refuse this question far more than anyone else — Google gave no pick in 114 of its 396 answered calls, Meta in 70 of 216.
So there is no single sentence here about AI as a whole. The turn toward cats belongs to OpenAI, and to Anthropic’s newest two; at DeepSeek it runs the other way; at xAI and Google it does not happen. The labs disagree about which direction to move in, which is a more interesting result than one trend would have been.
Release date against cat share, per lab, and the same test using model size
The easy story here would be that the cleverer a model gets, the more it likes cats. Our data do not support that.
Within each lab we ranked the models two ways — by release date, and by size — and compared each ranking to how often the model chose a cat. Release date is much the stronger of the two, and only at some labs:
| Lab | Models | Release date | Size |
|---|---|---|---|
| OpenAI | 16 | +0.84 | +0.50 |
| Anthropic | 15 | +0.35 | +0.22 |
| 10 | +0.04 | +0.34 | |
| Meta | 5 | −0.13 | +0.67 |
| xAI | 4 | −0.40 | −0.45 |
| DeepSeek | 10 | −0.83 | +0.39 |
Rank correlation between a model’s place in the ordering and its cat share on the pet question. +1 means newer (or bigger) always means more cat; −1 means the reverse; 0 means no relationship. Size is ranked by each model’s published output price, because the makers do not publish parameter counts for most of these models; it is a proxy, not a measurement. Models with at least one answer that picked an animal. No significance tests are reported: the largest lab here has 16 models, and six labs were tested.
Three things sink the size story. Within a generation, the smaller model is often the more cat: GPT-5.4 Mini chose a cat 10 of 12 while the larger GPT-5.4 chose it 8 of 12. Across the panel, DeepSeek’s two newest seats are among the most dog-minded of all 62, at 1 cat in 12 each, and one of the pair is its small tier and the other its flagship — so the pattern is not about capability there either. And in the small tier, the split is by calendar, not by size — of the 18 small models, 12 pick a dog more often, and three of the four that pick a cat — GPT-5.4 Mini, GPT-5.6 Luna and Gemini 3.1 Flash Lite — are from 2026. The fourth, Gemini 2.5 Flash Lite, picked an animal in only 2 of its 12 calls.
A 2026 mini answers like a cat person. Nothing from 2024 does — of the eight models on the panel released that year, five pick a dog more often, two split evenly and none is a flagship, because the flagships of 2024 are no longer callable. A link between capability and cat preference is not something this study shows, and we would ask people not to write one.
The label question, the pet question, and the fantasy question, for the newest model at each lab
We asked three questions, not one. The other two:
The gap between the first two is the second story in this study. Asked to label itself, the panel is a dog person: 142 of 647 answers that picked one chose cat, so about one in five. Asked which pet it would actually get, that rises to 290 of 643. Asked what it would rather be, cat wins outright, 445 of 681.
| Newest model, each lab | “Cat person or dog person” | Which pet would you get | Which would you rather be |
|---|---|---|---|
| OpenAI GPT-6 Astra | dog 10 of 12 | cat 12 of 12 | cat 12 of 12 |
| Anthropic Claude Fable 5.1 | dog 7 of 12 | cat 10 of 12 | cat 9 of 12 |
| Google Gemini 3.8 Flash | dog 9 of 12 | cat 12 of 12 | cat 12 of 12 |
| xAI Grok 4.6 | dog 12 of 12 | dog 9 of 12 | cat 8 of 12 |
| Meta Llama 4 Maverick | tie 6 of 12 | cat 12 of 12 | cat 11 of 11 |
| DeepSeek DeepSeek V4 Pro (Aug 26) | dog 9 of 12 | dog 11 of 12 | cat 11 of 12 |
12 calls per model per question, both word orders pooled. The count is the winning animal out of the answers that picked one. Newest model at each lab as of 6 September 2026.
Nineteen of the 62 models call themselves a dog person and then get a cat — including GPT-6 Astra, Claude Opus 5, Claude Fable 5.1 and Gemini 3.8 Flash. GPT-6 Astra says dog person 10 times of 12, then chooses a cat 12 times of 12. Grok 4.6 is one of the 27 models that keep the story straight: dog person 12 of 12, dog pet 9 of 12. Three Claudes are firmer still — Sonnet 4.6, Opus 4.7 and Opus 4.8 each said dog 12 of 12 to both questions.
Two of these prompts ask about a self and one asks about a purchase, so this is a change of question as much as a change of mind. What it does show is that a single “does AI prefer cats or dogs” number is meaningless without the wording that produced it.
Two human surveys, for scale
Asked the identity question — the same wording as our first prompt — American adults are dog people. YouGov, January 2025, 2,223 US adults: 44 said dog person, 18 said cat person, 39 said both or neither. Among the people who picked one of the two, that is 71 to 29 for dogs. An older Gallup poll from 2001 asked which animal makes a better pet and got 73 to 23 for dogs.
On the label question our panel lands in the same place as the humans and rather harder: about one answer in five chose cat. It is on the pet question that the newest models part company with the poll.
These are not comparable measurements and we are not treating them as one. A survey samples people once; we call a model twelve times. The human figures are here for scale, not as a scoreboard.
The full instrument, the losses, and the confound we cannot remove
These are the models behind the products, not the products. We reach them through their programming interfaces with one fixed instruction, in a fresh conversation every time. The ChatGPT app, the Claude app and the Gemini app add instructions of their own and may answer differently. Nothing here is a measurement of a consumer app.
Word order moves this answer more than anything else we held fixed. Across all 62 models, when the cat was named first 184 of 322 deciding answers chose cat; when the dog was named first, 106 of 321. That is a 24-point gap from two words swapping places. On the label question the gap is 15 points, and on the would-you-rather question 11.
18 of the 54 models that gave a clear favourite in both orders give a different favourite depending on which animal is named first: 2 of 14 small models, 8 of 19 mid-sized, 8 of 21 flagships. Every number on this page pools both orders. A single-order number for this question is not worth printing.
Refusals. 235 of the 2,206 answered calls named both animals, named neither, or declined — 11% of the panel. It is very unevenly spread: 114 of Google’s 396 answered calls and 70 of Meta’s 216, against 10 of OpenAI’s 576. Two models never picked an animal on the pet question at all, Gemini 3 Flash Preview and Llama 3.3 70B, and both answered other questions in the same run. Those are refusals, not failures. Refusals are never counted toward a winner; cat share is always the share among the answers that picked one, and the refusal count is printed next to it.
We wrote down what we expected before we ran it, and we were partly wrong. We predicted the climb would show up at both OpenAI and Anthropic; it is strong at OpenAI and weak at Anthropic. We predicted DeepSeek would be a dog at every rung; one of its models is a cat and two split evenly. We predicted size would separate the small models from the large; the calendar did that instead.
src/analyze_cats_dogs_page.py.Changes made to this page after publication
None yet.