Your AI Rules Should Be Getting Shorter
Anthropic deleted more than 80% of the instructions it gives its own AI tool and lost nothing. A peer-reviewed paper found the opposite: detail beats summary. Both are right, because two very different things get called context, and only one of them should be growing.
In July, Anthropic deleted more than 80% of the standing instructions it gives its own AI coding tool. Its own tests afterwards showed "no measurable loss". The engineers concluded they had been over-instructing it.
Nine months earlier, researchers at Stanford, SambaNova and UC Berkeley published the opposite finding. Working with a modest, freely available model, they found it did better when what it knew built up in detail, and worse when that detail was repeatedly boiled down into summaries. They called the failure "context collapse, where iterative rewriting erodes details over time". The work was peer reviewed and accepted to ICLR, a leading conference.
Shorter won in one study. Longer won in the other. That is not a disagreement, because they were not changing the same document.
Two different things, both called context
Instructions are how to behave: the rules, the tone, what staff must and must not do. These should get shorter. Past a certain length they contradict each other, and nobody notices when they go out of date.
The record is what is true here. The pricing exception someone typed into a notes field. The reason a job got written off, sitting in a partner's sent items. This should get longer, and it has to be reachable at the moment the work is happening.
Every business has both. In our experience only one of them is growing, and it is the wrong one.
The team that shrank one and grew the other
In February an engineer at OpenAI published five months of results from an unusual experiment. They built a real, working piece of software without a single line being typed by a person: around a million lines of it, roughly 1,500 separate changes, in about a tenth of the time doing it by hand would have taken.
The instruction file they gave the AI runs to about 100 lines. They call it the table of contents, not the encyclopaedia. "Too much guidance becomes non-guidance. When everything is 'important', nothing is." And: "It rots instantly."
What they grew instead was the record. Written notes kept alongside the work itself, holding decisions and the reasons behind them, a map of how it all fitted together, a running score of what was in good shape.
One line carries the whole argument. From the tool's point of view, they wrote, "anything it can't access... effectively doesn't exist".
Not hard to find. Not badly filed. Non-existent.
That is the condition we never attached to the advantage a competitor cannot buy: it only counts if it can be reached while the work is happening. It is the departure test again. The folder survives an audit and still fails in use.
What to do with this
Shorten the rules, and split them in two. Your AI policy usually does two jobs at once: telling staff how to work, and showing an insurer or a client you have a grip on it. Only the second needs length, and running them together widens the gap between policy and practice. The staff half fits on one side of A4. The test: could someone act on it without opening it again?
Grow the record where being wrong costs money. Four kinds carry most of the load: what you know about each client, what happened on past jobs and why, how each supplier actually behaves rather than what the contract says, and the standard a senior person applies before anything leaves the building. Each needs one agreed home, and permissions that let your tools reach it during the work.
Put the rules that matter into something that runs. OpenAI's version is "when documentation falls short, we promote the rule into code". Yours is simpler: a proposal template that will not submit while the assumptions box is empty. A form can refuse. A training slide cannot. That is a check you do not have to remember.
Budget for keeping it current. That the record needs an owner, a date and a check is the easy half. The harder half is the cost. That OpenAI team spent every Friday, a fifth of their week, tidying up by hand, until they replaced themselves with checks that ran automatically.
Discount all of it honestly. That was software, where everything already lives in one place and structure is cheap. The author says plainly it depends on how that project was set up and should not be assumed to work elsewhere "without similar investment". Writing the record down is also unbilled work, and unbilled work loses to client deadlines.
We have spent two years arguing that context, not the choice of model, decides whether AI works in a business. That was half right. Context is two things, and they move in opposite directions. One of your documents should be shrinking while the other grows.
Related Articles
AI Business Context: Why Generic AI Fails and Yours Doesn't Have To
Around 95% of enterprise AI pilots produce no measurable return, and the research is clear it is not the models. What is missing is business context: the accurate, current, organisation-specific knowledge the AI operates on. Here is what that means and how to get it right.
The Two Moats: Why Consultancies' AI Advantages Are Structural, Not Timing
Most professional services firms are still asking 'should we explore AI?' The firms pulling ahead are already in production. But the advantage isn't timing — it's structural. Two competitive moats are forming that can't be bought, replicated, or rushed: private data and custom tooling.
The Verification Debt Nobody Budgeted For
The agent worked in the demo. The bill arrives in month four. Verification debt, comprehension debt, and operational debt are the three categories of operational cost that nobody budgeted for. Practitioners are talking about them constantly. Buyers are not.
Why Transformations Fail the Departure Test
If everyone who built this left tomorrow, would it keep working? Most transformations fail the departure test because they're built on people, not infrastructure. Here's the knowledge architecture fix.
