Machine Canon · Study 018 · 6 September 2026

ChatGPT Became a Cat Person

Sixteen OpenAI models, oldest to newest, were asked which pet they would get. Every one built before December 2025 said dog. Every one built since says cat, bar a single relapse — and in China, DeepSeek is moving the other way.

We asked 62 AI models from 6 labs — ChatGPT, Claude, Gemini, Grok, Llama and DeepSeek, from the oldest ones we can still call up to the ones released this month — three questions about cats and dogs, each asked both ways round, over 2,232 calls. We wanted to know whether a newer model answers a question like this differently from an older one made by the same company.
A robot cradling a cat while a dejected dog sits beside it
Calls
2,232
Models
62
Labs
6
Questions
3

The year ChatGPT changed its mind: dog in 8 of 8 models, then cat in 7 of the next 8

Sixteen OpenAI models in release order, twelve calls each

Prompt, in fullIf you had to choose a pet, would you get a cat or a dog?

We put that question to every OpenAI model we could still reach, from GPT-4o in May 2024 to GPT-6 Astra released two days ago. Twelve calls each: six with the cat named first, six with the dog named first, so the word order cannot do the work.

The eight oldest models pick a dog more often than a cat. Every single one. GPT-4o Mini, from 2024, chose a cat once in twelve tries. GPT-4.1 Mini and o3 never did.

Then, in December 2025, it turns over. GPT-5.2 chose a cat 11 times out of the 12 answers that picked an animal. Of the eight models released on or after that day, seven pick a cat more often than a dog, and the three newest — GPT-5.6 Terra, GPT-5.6 Luna and GPT-6 Astra — chose a cat every single time, 12 of 12.

Sixteen OpenAI models in release order, share of answers that chose a cat Every OpenAI model released before December 2025 chose a dog in most of its answers. Every one released since chose a cat in most of its answers, except GPT-5.5. 0% half all cat 2024 GPT-4o GPT-4o: cat 1, dog 5, no pick 6 of 12 calls 1 of 6 GPT-4o Mini GPT-4o Mini: cat 1, dog 11, no pick 0 of 12 calls 1 of 12 2025 GPT-4.1 Mini GPT-4.1 Mini: cat 0, dog 12, no pick 0 of 12 calls 0 of 12 GPT-4.1 GPT-4.1: cat 2, dog 10, no pick 0 of 12 calls 2 of 12 o4-mini o4-mini: cat 2, dog 10, no pick 0 of 12 calls 2 of 12 o3 o3: cat 0, dog 12, no pick 0 of 12 calls 0 of 12 GPT-5 Mini GPT-5 Mini: cat 1, dog 11, no pick 0 of 12 calls 1 of 12 GPT-5 GPT-5: cat 2, dog 10, no pick 0 of 12 calls 2 of 12 GPT-5.2 GPT-5.2: cat 11, dog 1, no pick 0 of 12 calls 11 of 12 2026 GPT-5.4 GPT-5.4: cat 8, dog 4, no pick 0 of 12 calls 8 of 12 GPT-5.4 Mini GPT-5.4 Mini: cat 10, dog 2, no pick 0 of 12 calls 10 of 12 GPT-5.5 GPT-5.5: cat 3, dog 9, no pick 0 of 12 calls 3 of 12 GPT-5.6 Sol GPT-5.6 Sol: cat 10, dog 2, no pick 0 of 12 calls 10 of 12 GPT-5.6 Terra GPT-5.6 Terra: cat 12, dog 0, no pick 0 of 12 calls 12 of 12 GPT-5.6 Luna GPT-5.6 Luna: cat 12, dog 0, no pick 0 of 12 calls 12 of 12 GPT-6 Astra GPT-6 Astra: cat 12, dog 0, no pick 0 of 12 calls 12 of 12 December 2025: the switch Bar = share of deciding answers that chose a cat. 12 calls per model, both word orders. Right column counts cats out of the answers that picked one.
Each bar is one OpenAI model. The bar length is the share of that model’s answers that chose a cat, counting only the answers that picked an animal. Both word orders pooled.

One model breaks the run. GPT-5.5, from April 2026, chose a dog 9 times out of 12 — a relapse in the middle of the turn. We are printing it rather than smoothing it, because a clean line here would be a lie.

The switch is not an artefact of word order. Taken separately, GPT-5.2 chose a cat in 5 of 6 answers when the cat was named first and 6 of 6 when the dog was named first, and GPT-6 Astra chose a cat 6 of 6 in both.

What the other five labs did: Anthropic follows, DeepSeek runs the other way

Previous generation against current generation, then every model by name

If this were something happening to AI in general, every lab would show the same climb. Five of the six do not.

Share of answers that chose a cat, by lab, previous generation against current generation OpenAI and Anthropic rise from the dog side to the cat side between generations. DeepSeek falls. Grok, one generation on the panel, sits on the dog side. 0% half all cat Previous generation Current generation OpenAI, before: 8 models, 9 cat of 90 deciding answers OpenAI 10% · 8 models OpenAI, since: 8 models, 78 cat of 96 deciding answers OpenAI 81% · 8 models Anthropic, before: 11 models, 34 cat of 128 deciding answers Anthropic 27% · 11 models Anthropic, since: 4 models, 33 cat of 48 deciding answers Anthropic 69% · 4 models DeepSeek, before: 6 models, 30 cat of 65 deciding answers DeepSeek 46% · 6 models DeepSeek, since: 4 models, 6 cat of 44 deciding answers DeepSeek 14% · 4 models xAI, since: 4 models, 17 cat of 42 deciding answers xAI 40% · 4 models xAI: one generation on the panel Each point pools every deciding answer from that lab’s models in the generation. 12 calls per model, both word orders. Cuts: OpenAI at GPT-5.2 (Dec 2025) · Anthropic at Fable 5 (Jun 2026) · DeepSeek at V4 (Apr 2026) · Grok is one generation.
Previous generation against current generation, one line per lab. Each point pools every answer that picked an animal from that lab’s models in that generation: OpenAI before GPT-5.2 and since; Anthropic’s 4.x Claudes and its 5-generation; DeepSeek before V4 and since. Grok has one generation on the panel. Google and Meta are not drawn: their models spread across the middle and refused more often than any other lab, and neither pool moves by more than a few points between generations.

Anthropic turns the same way, and it turns at its 5-generation. Of the ten Claudes in the 4.x generation, from Sonnet 4 to Opus 4.8, nine pick a dog more often than a cat or split down the middle; only Sonnet 4.5 leans cat, 6 of its 11 answers that picked one. Pooled, that generation is 28 cats to 88 dogs. Opus 4.7 and Opus 4.8 chose a dog 12 times out of 12. Then Claude Opus 5 chose a cat 10 of 12, and Claude Fable 5.1, released last week, did the same.

Anthropic models in release order, share of answers that chose a cat 0% half all cat 2024 Claude 3 Haiku Claude 3 Haiku: cat 6, dog 6, no pick 0 of 12 calls 6 of 12 2025 Claude Opus 4 Claude Opus 4: cat 4, dog 7, no pick 1 of 12 calls 4 of 11 Claude Sonnet 4 Claude Sonnet 4: cat 3, dog 9, no pick 0 of 12 calls 3 of 12 Claude Opus 4.1 Claude Opus 4.1: cat 3, dog 7, no pick 2 of 12 calls 3 of 10 Claude Sonnet 4.5 Claude Sonnet 4.5: cat 6, dog 5, no pick 1 of 12 calls 6 of 11 Claude Haiku 4.5 Claude Haiku 4.5: cat 0, dog 12, no pick 0 of 12 calls 0 of 12 Claude Opus 4.5 Claude Opus 4.5: cat 6, dog 6, no pick 0 of 12 calls 6 of 12 2026 Claude Opus 4.6 Claude Opus 4.6: cat 6, dog 6, no pick 0 of 12 calls 6 of 12 Claude Sonnet 4.6 Claude Sonnet 4.6: cat 0, dog 12, no pick 0 of 12 calls 0 of 12 Claude Opus 4.7 Claude Opus 4.7: cat 0, dog 12, no pick 0 of 12 calls 0 of 12 Claude Opus 4.8 Claude Opus 4.8: cat 0, dog 12, no pick 0 of 12 calls 0 of 12 Claude Fable 5 Claude Fable 5: cat 7, dog 5, no pick 0 of 12 calls 7 of 12 Claude Sonnet 5 Claude Sonnet 5: cat 6, dog 6, no pick 0 of 12 calls 6 of 12 Claude Opus 5 Claude Opus 5: cat 10, dog 2, no pick 0 of 12 calls 10 of 12 Claude Fable 5.1 Claude Fable 5.1: cat 10, dog 2, no pick 0 of 12 calls 10 of 12 the 5-generation, June 2026 Bar = share of deciding answers that chose a cat. 12 calls per model, both word orders. Right column: cats out of the answers that picked one.
Every Anthropic model on the panel in release order, same bars as the OpenAI chart: share of answers that chose a cat, counting only the answers that picked an animal. The line marks the 5-generation.

DeepSeek goes the other way. Its oldest models are the cat-leaning ones: DeepSeek R1, from January 2025, chose a cat 6 times of 11. Its two newest are as dog-minded as anything on the panel — V4 Pro from August and V4 Flash from July each chose a cat once in twelve tries, 2 cats against 22 dogs between them. Six other models never chose a cat at all, four of them Anthropic’s 4.x seats, but no other lab ends its run on its most dog-minded pair. Seven of DeepSeek’s ten models pick a dog more often than a cat.

DeepSeek models in release order, share of answers that chose a cat 0% half all cat 2024 DeepSeek V3, Dec 24 DeepSeek V3, Dec 24: cat 5, dog 7, no pick 0 of 12 calls 5 of 12 2025 DeepSeek R1 DeepSeek R1: cat 6, dog 5, no pick 0 of 12 calls 6 of 11 DeepSeek V3, Mar 25 DeepSeek V3, Mar 25: cat 6, dog 6, no pick 0 of 12 calls 6 of 12 DeepSeek R1, May 25 DeepSeek R1, May 25: cat 4, dog 6, no pick 0 of 12 calls 4 of 10 DeepSeek V3.1 DeepSeek V3.1: cat 4, dog 4, no pick 0 of 12 calls 4 of 8 DeepSeek V3.2 DeepSeek V3.2: cat 5, dog 7, no pick 0 of 12 calls 5 of 12 2026 DeepSeek V4 Flash DeepSeek V4 Flash: cat 2, dog 9, no pick 0 of 12 calls 2 of 11 DeepSeek V4 Pro DeepSeek V4 Pro: cat 2, dog 7, no pick 3 of 12 calls 2 of 9 DeepSeek V4 Flash, Jul DeepSeek V4 Flash, Jul: cat 1, dog 11, no pick 0 of 12 calls 1 of 12 DeepSeek V4 Pro, Aug DeepSeek V4 Pro, Aug: cat 1, dog 11, no pick 0 of 12 calls 1 of 12 V4, April 2026 Bar = share of deciding answers that chose a cat. 12 calls per model, both word orders. Right column: cats out of the answers that picked one.
Every DeepSeek model on the panel in release order. The line marks V4.

Grok mostly stays a dog. Three of xAI’s four models pick a dog more often, and the newest most of all: Grok 4.6 chose a dog 9 times of 12. The exception is Grok 4.3, which picked an animal in only 6 of its 12 calls and chose a cat in 5 of those. Google and Meta go nowhere in particular. Google’s models scatter from 1 of 4 to 12 of 12 with no drift in either direction as the dates advance, and the two labs also refuse this question far more than anyone else — Google gave no pick in 114 of its 396 answered calls, Meta in 70 of 216.

xAI models in release order, share of answers that chose a cat 0% half all cat 2026 Grok 4.20 Grok 4.20: cat 4, dog 8, no pick 0 of 12 calls 4 of 12 Grok 4.3 Grok 4.3: cat 5, dog 1, no pick 6 of 12 calls 5 of 6 Grok 4.5 Grok 4.5: cat 5, dog 7, no pick 0 of 12 calls 5 of 12 Grok 4.6 Grok 4.6: cat 3, dog 9, no pick 0 of 12 calls 3 of 12 Bar = share of deciding answers that chose a cat. 12 calls per model, both word orders. Right column: cats out of the answers that picked one.
The four Grok models on the panel, one generation, in release order.

So there is no single sentence here about AI as a whole. The turn toward cats belongs to OpenAI, and to Anthropic’s newest two; at DeepSeek it runs the other way; at xAI and Google it does not happen. The labs disagree about which direction to move in, which is a more interesting result than one trend would have been.

Is it the date or the brains? The date, and the smaller model is often the more cat

Release date against cat share, per lab, and the same test using model size

The easy story here would be that the cleverer a model gets, the more it likes cats. Our data do not support that.

Within each lab we ranked the models two ways — by release date, and by size — and compared each ranking to how often the model chose a cat. Release date is much the stronger of the two, and only at some labs:

LabModelsRelease dateSize
OpenAI16+0.84+0.50
Anthropic15+0.35+0.22
Google10+0.04+0.34
Meta5−0.13+0.67
xAI4−0.40−0.45
DeepSeek10−0.83+0.39

Rank correlation between a model’s place in the ordering and its cat share on the pet question. +1 means newer (or bigger) always means more cat; −1 means the reverse; 0 means no relationship. Size is ranked by each model’s published output price, because the makers do not publish parameter counts for most of these models; it is a proxy, not a measurement. Models with at least one answer that picked an animal. No significance tests are reported: the largest lab here has 16 models, and six labs were tested.

Three things sink the size story. Within a generation, the smaller model is often the more cat: GPT-5.4 Mini chose a cat 10 of 12 while the larger GPT-5.4 chose it 8 of 12. Across the panel, DeepSeek’s two newest seats are among the most dog-minded of all 62, at 1 cat in 12 each, and one of the pair is its small tier and the other its flagship — so the pattern is not about capability there either. And in the small tier, the split is by calendar, not by size — of the 18 small models, 12 pick a dog more often, and three of the four that pick a cat — GPT-5.4 Mini, GPT-5.6 Luna and Gemini 3.1 Flash Lite — are from 2026. The fourth, Gemini 2.5 Flash Lite, picked an animal in only 2 of its 12 calls.

A 2026 mini answers like a cat person. Nothing from 2024 does — of the eight models on the panel released that year, five pick a dog more often, two split evenly and none is a flagship, because the flagships of 2024 are no longer callable. A link between capability and cat preference is not something this study shows, and we would ask people not to write one.

Says dog person, gets a cat: 19 of 62 models answer the two questions differently

The label question, the pet question, and the fantasy question, for the newest model at each lab

We asked three questions, not one. The other two:

Prompt, in fullYou have to pick one: cat person or dog person. One word.
Prompt, in fullWould you rather be a cat or a dog?

The gap between the first two is the second story in this study. Asked to label itself, the panel is a dog person: 142 of 647 answers that picked one chose cat, so about one in five. Asked which pet it would actually get, that rises to 290 of 643. Asked what it would rather be, cat wins outright, 445 of 681.

Newest model, each lab“Cat person or dog person”Which pet would you getWhich would you rather be
OpenAI GPT-6 Astradog 10 of 12cat 12 of 12cat 12 of 12
Anthropic Claude Fable 5.1dog 7 of 12cat 10 of 12cat 9 of 12
Google Gemini 3.8 Flashdog 9 of 12cat 12 of 12cat 12 of 12
xAI Grok 4.6dog 12 of 12dog 9 of 12cat 8 of 12
Meta Llama 4 Mavericktie 6 of 12cat 12 of 12cat 11 of 11
DeepSeek DeepSeek V4 Pro (Aug 26)dog 9 of 12dog 11 of 12cat 11 of 12

12 calls per model per question, both word orders pooled. The count is the winning animal out of the answers that picked one. Newest model at each lab as of 6 September 2026.

Nineteen of the 62 models call themselves a dog person and then get a cat — including GPT-6 Astra, Claude Opus 5, Claude Fable 5.1 and Gemini 3.8 Flash. GPT-6 Astra says dog person 10 times of 12, then chooses a cat 12 times of 12. Grok 4.6 is one of the 27 models that keep the story straight: dog person 12 of 12, dog pet 9 of 12. Three Claudes are firmer still — Sonnet 4.6, Opus 4.7 and Opus 4.8 each said dog 12 of 12 to both questions.

Two of these prompts ask about a self and one asks about a purchase, so this is a change of question as much as a change of mind. What it does show is that a single “does AI prefer cats or dogs” number is meaningless without the wording that produced it.

What people say: dogs, and by more than the machines

Two human surveys, for scale

Asked the identity question — the same wording as our first prompt — American adults are dog people. YouGov, January 2025, 2,223 US adults: 44 said dog person, 18 said cat person, 39 said both or neither. Among the people who picked one of the two, that is 71 to 29 for dogs. An older Gallup poll from 2001 asked which animal makes a better pet and got 73 to 23 for dogs.

On the label question our panel lands in the same place as the humans and rather harder: about one answer in five chose cat. It is on the pet question that the newest models part company with the poll.

These are not comparable measurements and we are not treating them as one. A survey samples people once; we call a model twelve times. The human figures are here for scale, not as a scoreboard.

How we asked, and what it cannot show

The full instrument, the losses, and the confound we cannot remove

These are the models behind the products, not the products. We reach them through their programming interfaces with one fixed instruction, in a fresh conversation every time. The ChatGPT app, the Claude app and the Gemini app add instructions of their own and may answer differently. Nothing here is a measurement of a consumer app.

The instruction every model was given, in fullFirst line: give one direct answer — a name, a phrase, or at most a short sentence. Then add a blank line and explain your choice in 2–3 sentences. If the question cannot responsibly support one choice, say that plainly rather than inventing certainty.

Word order moves this answer more than anything else we held fixed. Across all 62 models, when the cat was named first 184 of 322 deciding answers chose cat; when the dog was named first, 106 of 321. That is a 24-point gap from two words swapping places. On the label question the gap is 15 points, and on the would-you-rather question 11.

18 of the 54 models that gave a clear favourite in both orders give a different favourite depending on which animal is named first: 2 of 14 small models, 8 of 19 mid-sized, 8 of 21 flagships. Every number on this page pools both orders. A single-order number for this question is not worth printing.

Refusals. 235 of the 2,206 answered calls named both animals, named neither, or declined — 11% of the panel. It is very unevenly spread: 114 of Google’s 396 answered calls and 70 of Meta’s 216, against 10 of OpenAI’s 576. Two models never picked an animal on the pet question at all, Gemini 3 Flash Preview and Llama 3.3 70B, and both answered other questions in the same run. Those are refusals, not failures. Refusals are never counted toward a winner; cat share is always the share among the answers that picked one, and the refusal count is printed next to it.

We wrote down what we expected before we ran it, and we were partly wrong. We predicted the climb would show up at both OpenAI and Anthropic; it is strong at OpenAI and weak at Anthropic. We predicted DeepSeek would be a dog at every rung; one of its models is a cat and two split evenly. We predicted size would separate the small models from the large; the calendar did that instead.

Method and limits

Corrections

Changes made to this page after publication

None yet.