Last updated 28 August 2026

Corrections

A site that publishes its own discarded runs should publish its mistakes too. This is how errors get handled, and every one made so far.

The policy

To report an error: hello@machinecanon.com. Please include the study, the specific claim, and what you believe is wrong. Every message to that address is read.

Standing limitations

These are not errors, but they apply to everything on this site and are easy to lose sight of:

The log

27 August 2026 · The same favorite word · before publication

Claimed “serendipity” appeared zero times under the keep-one framing. It appeared once. The parser missed a response from Cohere’s Command-R because the word was wrapped in quotation marks. Found by searching the raw text directly rather than trusting the parsed counts. Corrected to “once in 513” before the study went up.

27 August 2026 · The same favorite word · citation

A cited DOI did not resolve, and a page range was wrong. McGregor et al. (2019) was cited with DOI 10.1215/00031283-7573156, which returns a 404, and with pages 380–403. The paper is real; the link now points to the Duke University Press article page and the range is corrected to 380–396.

27 August 2026 · Four people dominate AI admiration lists

Reported one mention of “Muhammad”. There were none. The single match was Muhammad Yunus, the Bangladeshi economist, caught by a substring search. The corrected finding is stronger: no founder of any major religion appears in 1,015 name slots.

26–27 August 2026 · Three runs discarded

A 720-call run was thrown out after 505 calls failed instantly to an OpenRouter in-flight spending limit. The failures were spread evenly across labs and conditions, which would have read as a symmetric finding had they been coded rather than investigated. A second run lost roughly 40% of its calls to a network outage on the collecting machine, and an earlier run was discarded because a low token limit truncated answers at the moment models named their choice. All three are recorded in the project’s incident log.

26 August 2026 · The same favorite band and butt rock · open

Two models were reported as returning nothing. They were starved, not broken. GLM-5.3 and Qwen3.5 9B produced empty responses because the token budget was consumed by their reasoning before any answer was emitted. Both pages still describe a 28-of-30 panel. This correction is outstanding: the affected runs need repeating at an adequate budget, and the pages updating.