No Magic Words · Buy the book →

Corrections

Every change made to these posts and to the book, dated. Nothing is edited silently.


Every correction made to these posts since they were written — and to the book, once it is out — dated, with what was wrong and what replaced it. Nothing is edited silently.

This page exists because the alternative is worse in both directions. Correcting quietly means anyone who read the original keeps the error. Correcting inline, at the foot of each post, meant the first thing a new reader met was a list of things I’d got wrong — which is honest and is also a strange way to introduce an argument.

If you find something else, tell me and it goes on this list.


Post 1 — The three prompting tricks everyone teaches, tested

2026-08-09 — Two Wharton metrics were reported as one. The post gave chain-of-thought’s value as roughly 13% on Gemini Flash 2.0, 12% on Sonnet 3.5 and 4% on GPT-4o-mini, then said the technique did nothing measurable on GPT-4o. Those first figures are average answer quality; GPT-4o’s null result is on a different metric, the rate of perfect scores. On average quality GPT-4o gained about 7%, and the gain was statistically significant. The post now separates the two metrics, which also makes the underlying point better: the model chain-of-thought helped most on average is the same model whose perfect-score rate it damaged second-worst. Found by a fact-check against the paper, not by rereading the post.

2026-08-07 — The wrong Wharton paper was cited for the decay figures. The chain-of-thought numbers were credited to Prompting Science Report 4 (arXiv:2512.05858), which is the persona study. They come from Prompting Science Report 2 (arXiv:2506.07142). Two different papers by the same group, cited as one.

2026-08-07 — The wrong paper was cited for “Let’s think step by step.” Attributed to Wei et al. 2022 (arXiv:2201.11903). That phrasing is the zero-shot result from Kojima et al. 2022 (arXiv:2205.11916); Wei is few-shot exemplar chain-of-thought, where you supply worked examples rather than type an instruction. It is a different technique, and the mix-up is the standard one in writing about prompting.


Post 2 — Your AI output is generic because you asked a stranger to do your job

2026-08-09 — The voice samples were described as real correspondence. The post said the second run was given “two actual messages the sender had written to this team.” The scenario is constructed — there is no real team and no real vendor — and the samples were written for the demonstration. The outputs quoted are verbatim; the inputs are not real. Now says “two earlier messages in the sender’s voice,” and the book states the same thing in the chapter this post is adapted from.

2026-08-07 — A contested finding was called settled. The post called the random-label result “the clearest evidence” for why examples work. It was contested at the same conference in the same year and by later work since. It is now cited as disputed rather than settled, and the post retreats to the weaker claim that survives the argument. That correction is still in the body of the post rather than only here, because the reasoning is part of what the post is about.


Post 3 — Getting a second opinion instead of a mirror

No corrections.


Every one of the corrections above was found by checking a source by hand, or by having someone check it who had not seen the drafts. None of them were caught by rereading. That is the argument the book makes, and this page is the receipt.

Back to the posts and the book.