No Magic Words

The three prompting tricks everyone teaches, tested


If you’ve read anything about prompting, you’ve been told some version of three things: assign the model a role (“you are an expert marketing strategist”), be polite to it, and tell it to think step by step.

They appear in nearly every prompt guide, paid course, and 300-prompt ebook in circulation.

All three have been tested. Here’s what the testing found.


Personas do nothing

A 2024 study ran 162 different role assignments across 2,410 factual questions and four model families. Not a blog experiment — a serious piece of work.

Adding a persona produced no consistent improvement. The average effect was slightly negative.

A team at Wharton replicated the question in 2025 on harder benchmarks and found expert personas had no significant effect on accuracy at all. Off-domain personas produced marginal differences and sometimes degraded performance; the clearer negative result was for low-knowledge personas, which often reduced accuracy.

So the realistic range of this technique runs from “does nothing” to “occasionally costs you something.”

Telling a model it’s a senior strategist doesn’t make it reason like one. It changes the vocabulary it decorates the answer with — which is exactly the kind of change that feels like an improvement and isn’t.

Politeness barely matters

Researchers tested whether courteous phrasing improved output. On weaker, older models it had some effect. On GPT-4-class models, little.

Be polite if you want to. There are decent human reasons to keep the habit — you probably don’t want to practice being curt for several hours a day. Just don’t do it believing it buys better work.

“Think step by step” is now discouraged by the vendors

This is the one that should give you pause, because it was, for about two years, the single most-recommended prompting technique in the world.

And here’s the part that matters: it was excellent advice. It came from a legitimate 2022 research finding — telling a model to reason step by step substantially improved performance on hard problems. The paper was real. The effect was real. Anyone teaching it was teaching something true.

It’s now obsolete. OpenAI’s own guidance for its reasoning models says plainly that because these models reason internally, prompting them to think step by step is unnecessary — and notes that on at least one model, the technique makes performance worse.

Wharton put numbers on the decay: still worth 11–13% on older non-reasoning models, but around 3% on current reasoning models, often not worth the added latency.


The pattern underneath

Nothing went wrong with “think step by step.” The technique didn’t fail. It got absorbed.

What was once an external trick you had to know became a behavior trained into the model itself. The prompt became unnecessary because the principle won.

That’s the pattern worth internalizing: the trick expired; the principle didn’t.

Anyone who learned “type these words” has to relearn. Anyone who understood why it worked — that hard problems benefit from being broken into steps — needed to change nothing at all. They just stopped typing a sentence the model now does on its own.

Tricks exploit the gap between what a model does and what you need. As models improve, those gaps close, and the tricks die with them.

Now consider what that implies about a book selling you three hundred prompts.


What hasn’t expired

Four things have been true since 2022 and are still true through several complete generations of model capability:

None of those are exciting. They’re also the only advice from 2022 that survived intact.


One honest complication

The tidy version of this argument would be “phrasing doesn’t matter.” That’s not what the evidence says.

Researchers have shown that purely cosmetic changes to a prompt’s format — ones that don’t alter meaning at all — can swing accuracy dramatically, in one study by up to 76 points. And the paper is explicit that this brittleness is not mitigated by model size. It hasn’t been going away as models improve.

So wording still matters. But notice what that argues for. If improvised phrasing has large effects nobody can predict or intuit, the answer isn’t to improvise better — there’s no skill to acquire, because the effects don’t follow any learnable pattern.

The answer is to stop improvising. Build the input once, in a form you’ve seen work, and reuse it.

Brittleness is an argument for structure, not for cleverness.


This is adapted from Chapter 1 of NO MAGIC WORDS — how to get real work out of AI, and keep getting it when the models change.

Two more posts in this series: why context beats phrasing (with a full before/after demonstration), and how to get a second opinion from AI instead of a mirror.

Sources: Zheng et al. 2024 (arXiv:2311.10054) · Wharton GenAI Labs 2025 (arXiv:2512.05858) · Yin et al. 2024 (arXiv:2402.14531) · OpenAI reasoning best-practices guide · Wei et al. 2022 (arXiv:2201.11903) · Sclar et al. 2024 (ICLR, arXiv:2310.11324)

Get the context pack template

The fifteen-minute setup from the book's appendix — the single change that improves most people's output today. Plus the rest of this series, and a note when the book is out. No other email.