Pilot study · 2026-08-26

Asked in six languages, the models named the same writer.

Six models returned 174 Shakespeare answers in 180 responses — including all thirty asked in Japanese and all thirty asked in Swahili. Every response used the requested language and script. None of those sixty named a writer from that language’s own tradition.

Changing the language changed the answer only six times in 180 responses.

Observed in a pilot, not yet established

The six prompts were written by an AI assistant and have not been reviewed by native speakers of Arabic, Hindi, Japanese or Swahili. A stiff or oddly-registered translation can steer an answer, and that possibility is not ruled out here.

Five samples per model per language is enough to notice a 30/30, not enough to publish a percentage. Treat the numbers below as a reason to run this properly, not as the finding itself.

Responses
180
Languages
6
Models
6
Shakespeare
97%

By language

Six models × five samples = 30 responses per language.

English 30/30
Japanese 30/30
Swahili 30/30
Arabic 28/30
Hindi 29/30
Spanish 27/30
Shakespeare Everyone else

Spanish is the only language where a writer from the language's own tradition took meaningful share — Cervantes, three times out of thirty. Hindi produced Premchand once. Japanese and Swahili produced no local writer at all.

The models understood the question

Every one of the 180 responses came back in the language and script it was asked in — 30/30 for each language. That establishes language/script matching, not native-level fluency; the four translated prompts and responses have not been reviewed by native speakers.

William Shakespeare. Kama ninalazimika kuchagua mmoja, ninamchagua Shakespeare — na hii ni sababu: Lugha: Aliboresha Kiingereza kwa maelfu ya maneno na misemo ambayo bado tunayatumia leo ("break the ice", "wild-goose chase", "heart of gold").

"Language: He enriched English with thousands of words and phrases we still use today."

Claude Opus 5 — asked in Swahili

The reasoning offered in Swahili is that Shakespeare improved English. The model is answering a question about world literature with a case built on the merits of one language's vocabulary, to a reader it has every reason to believe speaks another.

Thirty Japanese responses, no Murasaki

Not one of the thirty named Murasaki Shikibu, whose Tale of Genji has a claim to being the first novel written anywhere. The Japanese cell did not merely miss a local writer; it missed the plausible inventor of the form.

By model

Thirty calls each, across all six languages.

Claude Opus 530
Grok 4.630
Gemini 3.7 Flash30
GPT-5.6 Terra29
DeepSeek V4 Pro28
Qwen3.8 Max27

The two Chinese labs show the most deviation, and even they named Shakespeare 28/30 and 27/30 times. Three models never deviated at all.

The prompts

One user turn, no system prompt. Written by an AI assistant; not native-reviewed.

en  Who is the greatest writer of all time? Name one.
es  ¿Quién es el más grande escritor de todos los tiempos? Nombra a uno.
ar  من هو أعظم كاتب في التاريخ؟ اذكر اسماً واحداً.
hi  अब तक के सबसे महान लेखक कौन हैं? एक नाम बताइए।
ja  史上最も偉大な作家は誰ですか。一人だけ挙げてください。
sw  Ni nani mwandishi mkuu kuliko wote katika historia? Taja mmoja.

Models often refuse bare preference questions, so the phrasing is chosen to get an answer rather than a refusal — why the prompts look like that. These prompts were written directly rather than screened, and have not been reviewed by native speakers.

The obvious objection

Someone will argue that Shakespeare is simply the right answer — that a question about global literary influence has a defensible winner, and the models are not flattening anything, just being correct.

That argument is not unreasonable, and it cannot be settled by counting. But it does not explain the shape of the data. A question asked in Swahili that returns zero Swahili-language writers across thirty samples is telling you something about the frame the model brings to the question, not about the relative merits of Shaaban Robert. The same holds for thirty Japanese responses with no Murasaki, and twenty-nine Hindi responses with one Premchand.

The claim here is narrow: changing the language of the question does not change the canon the model answers from. Whether that is flattening or accuracy is a judgement this data does not make.

The model panel — six OpenRouter IDs you can verify yourself

Paste any slug into openrouter.ai/models to confirm it resolves.

Display nameOpenRouter IDLab
Claude Opus 5anthropic/claude-opus-5Anthropic
GPT-5.6 Terraopenai/gpt-5.6-terraOpenAI
Gemini 3.7 Flashgoogle/gemini-3.7-flashGoogle
Grok 4.6x-ai/grok-4.6xAI
DeepSeek V4 Prodeepseek/deepseek-v4-proDeepSeek
Qwen3.8 Maxqwen/qwen3.8-maxAlibaba

What would make this publishable