I build software for small businesses. About a dozen clients, no employees, and roughly seventy agents on my squad, which is a ratio of seventy to one. I put the whole thing together myself over about two years. In the middle of this year that system started refusing to let me work, politely, with citations, for reasons that were wrong. It took a month to understand and two days to clean up, and that ratio is the part worth carrying: the diagnosis is the expensive half, and the repair is smaller than the dread that surrounds it. What I learned is not specific to my setup. If your agents read prose that a human wrote, you have this, whether or not it has bitten you yet.
You are dreading the session
You sit down, and before anything happens you already know how it will go. The agent will get partway into the work and stop, for a reason that sounds serious, and you will spend the next forty minutes finding out whether the reason is true.
Nothing is broken, which is the part that is hard to explain. The logs are clean. Nothing errored. The agent did not invent anything or lose the thread. It read a page, quoted a sentence, drew the obvious conclusion, and handed you the file. The reasoning was sound. A careful person reading the same page would have said the same thing.
And you cannot adjudicate it, because adjudicating means going and checking, and checking means an afternoon you do not have.
So you make a call, the way a reasonable person under time pressure makes a call. Sometimes you override it and the day continues. Sometimes you believe it and the work does not get done. Either way something gets spent, and it is not documentation quality and it is not model reliability. It is your own willingness to take the next alarm seriously.
After enough of these, you stop feeling alarm at all. What you feel instead is a specific tiredness: being stopped again, by something that will take an afternoon to unpick, at the end of a day, over a task you have already postponed more times than you want to count.
And there is probably a piece of work sitting in your list that you keep not doing. Every week it has a plausible reason attached, and the reason is always real, and the more urgent thing really is more urgent. Which is how a person spends a month avoiding half an hour of work without once catching themselves avoiding anything.
If none of that is familiar, you can go, and I mean that kindly. If it is familiar, the rest is about why it happens, why it is not a failure of yours, and what held up in my house.
The problem has three floors
On the surface, the problem is false stops. Your agent raises a flag, you investigate, and the flag turns out to be attached to a sentence that used to be true. Nothing was wrong with the flag. It was about a world that has moved.
Underneath that is the part nobody says out loud. You no longer trust the system you built, and you cannot say why. There is no bug to file and no incident to write up. There is only a slow shift in how you feel about opening the work, and a quiet, unpleasant suspicion that the problem is you. That you were not disciplined enough with your notes. That somebody more organized would not be here.
And underneath that is the thing that actually matters. A safety system you cannot trust is worse than none, because it costs the same and protects less. If you learn to ignore your agents, the safety rule dies, since a gate nobody honors is not a gate. If you obey every agent, the product dies, because you never ship. Both jaws close at once, and there is no correct place to stand between them. It should not be this way, and it is not a law of nature that it is. It is a consequence of an arrangement, and arrangements can be changed.
It is easy to read a story about one person's repository as a story about repositories, so let me be clear about who has it.
You have it if you attach project knowledge to an assistant and expect it to be honored, or run a custom assistant with a folder of reference documents behind it, or keep an instructions file in the root of a repository that every session reads. You have it if you keep a folder of rules, or a memory server your agent queries, or an internal wiki, or a directory of runbooks that somebody wrote carefully on a good day. The stack does not matter, and neither does the name it goes by: memory, context, knowledge, standards.
The only requirement is that prose carries authority in your setup, and a human maintains that prose by hand. If both of those are true, the failure I am about to describe is available to you, and the better your system is at honoring what you wrote, the more efficiently it will deliver the parts that have gone off.
It is not your fault, and it is not your agent's
Here is the part I would most like you to keep.
The alarm was faithful. The agent read what it was given and reasoned correctly from it. There is no deception anywhere in the chain, at any layer, by anybody. The model did not hallucinate, retrieval did not fail, nothing was embellished. Given the sentence it had, the conclusion followed, and it was wrong only because the sentence had stopped matching the world.
And nobody failed to update the documents.
That one took me a long time to accept, because it removes the villain. The natural story about a stale document is that somebody got lazy, and under that story the fix is conscientiousness: write it down when you change it, do a documentation pass at the end of the sprint, put a reminder on the calendar. All of it rests on one picture of the world. A human wrote the page, a human changed the system, and the same human failed to close the loop. That picture held until about eighteen months ago. It does not describe the world you are working in now.
A good day of agentic coding produces thirty documents. A very good day produces a hundred. Not because anybody decided documentation was important, and not because anybody is being verbose. Writing documents is how the work now gets done. A plan comes out of a planning phase, resume notes out of an implementation phase, findings out of a review. Hundreds of thousands of words, in a day, often at a better standard than a tired human would manage at five in the afternoon.
And every one of those documents is accurate at the moment it is created. They are not sloppy and they are not made up. They describe, correctly, the state of the world in the hour they were written.
So the arithmetic that matters is not effort against laziness. It is production rate against verification rate. A person can properly verify a handful of claims an hour, where verifying means going and looking: running the command, reading the deployed configuration, checking that the sentence still corresponds to something outside it. Call it ten an hour with nothing interrupting you. A day of agentic work emits several thousand. Triple your conscientiousness and nothing changes, because tripling a number that is three orders of magnitude too small leaves it three orders of magnitude too small.
Conscientiousness was never the variable. You did not get sloppy. You got productive, and productivity has a byproduct, and nobody built the ventilation, because until recently nobody needed any.
Here is what it is called
Two names, given early, because a problem you cannot name is one you cannot raise with anyone, including yourself.
The condition in the corpus is semantic exhaust: the accumulation of authority-shaped prose that was accurate when it was written and has stopped being true.
I chose exhaust deliberately. Not pollution, which implies somebody was dumping. Not rot, which implies neglect. Not debt, which implies a choice to defer, and there was no moment of deferral to point at. Exhaust is what an engine emits while doing useful work. Nobody is at fault for it, and the answer to it is ventilation rather than shame.
Both halves carry weight. Authority-shaped means the prose has institutional force, because a person wrote it down as a rule or because the document claims to outrank other documents. Accurate when written means the sentence was true on the day. Not a guess, not an invention. True, and then quietly overtaken by the world.
The second name is for what it does to you, and it is not mine. Alarm fatigue comes out of clinical medicine, where it has been studied for years.
The shape is this. Put a patient on monitors. The monitors alarm. Most of the alarms do not indicate anything requiring action, because the thresholds are conservative, the sensors are noisy, and a patient rolling over will set one off. Clinicians hear a great many in a shift, and they learn, correctly, from real evidence, that the alarm carries very little information. Response times stretch. Volumes get turned down. And then one of the alarms is real, and it gets the response the false ones earned.
People have been hurt by that.
The literature exists, and hospitals have spent serious money and serious attention on it. I am not going to hand you a statistic, because I have not done that reading closely enough to quote it, and quoting a number I have not checked would be a strange thing to do in a piece about sentences that stopped being true. Go and look it up. The mechanism is what I am borrowing, and the mechanism is not in dispute. So: semantic exhaust is what is in the corpus. Alarm fatigue is what the corpus does to the person. The claim I am actually making, and the one to hold me to, is the arrow between them. A corpus full of sentences that were true once, faithfully retrieved and faithfully defended, produces in a human being the exact clinical pattern, and the human ends up unable to hear the alarm that matters.
How I know
There was a test I could not make myself walk toward. Take a booking through a preview environment, click the buttons a customer would click, watch the record land where it should. Half an hour of work, and it sat on my list for most of a month.
When I finally sat down to it, an agent stopped me. It had found a document and quoted it, and the quoted sentence said the deployed service held no payment credentials at all, so no payment path could move money, so the test could prove nothing. Given that sentence, the conclusion followed.
So I did the thing I should have done in July, and it took four seconds. I ran a read-only command that asks a deployed service to list the names of the secrets it holds. Names only, never values.
Five names came back. Four were payment keys, including the live pair. The fifth was an encryption key that protects other secrets at rest. The documents that had been stopping me for a month said the list was empty, and it had not been empty for a while.
Then the part that made me put my hands flat on the desk. A client had already been taking card payments at a physical counter for ten days. Real money, real customers, a shop I could drive to. Money does not move through a counter without live payment credentials behind it. The proof that the documents were wrong had been ringing up sales in front of me since the first of August, and I walked past it every day, because it lived in the world and the claim lived in a file, and I was checking the file.
There is an unflattering part. In February I built a detector for exactly this, and found it later on a backup drive, dated the second of February, complete and functional, with a report header reading top ten files most likely to be lying. I had been telling people I identified the problem and never built the thing. I built it, shelved it, and forgot I had, because I only knew the problem in theory, and a problem you have only reasoned your way to cannot compete with the things that are actually on fire. It has never burned you.
The reason this is worth your attention is that I have both arms of the experiment. Most writing about documentation discipline has one: an opinion and a tidy repository. I have the cost of doing nothing, which was a product that would not ship for a month. And I have the cost of doing the work: two days of unglamorous cleanup, a framework I had to build twice, and an audit on my own cleanup that came back with the verdict false. I would not have volunteered for that experiment, and almost nobody else has run it.
Why keeping the documents updated is not the answer
The instinct, once you see the problem, is to promise to do better. Update the pages when the world changes. Add a review step. Be more disciplined.
A promise to try harder is not a structure. It asks one finite person for more of a thing they already do as much of as they can, against a production rate three orders of magnitude past them. It holds for two weeks, then a client emergency lands and the plan quietly fails, and you feel like you failed, which is the worst part of it.
The second reason is architectural rather than motivational. Grep is not search. When an agent goes looking for the rule about some topic, what comes back is not a document. What comes back is a chunk: a window of text around a keyword hit, some lines above and below, chosen by a mechanism that knows nothing about meaning. The unit of consumption is a fragment.
Nobody has ever written a page expecting it to be read that way. You wrote a title, a date near the top, an opening paragraph framing the whole thing as a proposal, and then, under a heading saying proposed approach, a confident sentence about how production would work. Every framing device you supplied sits outside the window. The chunk holds the confident sentence and perhaps four lines around it.
The veto is never in the chunk. Not the date. Not the heading that said proposed. Not the paragraph at the top saying none of it had been built yet. And the agent does not know the material is missing, because from inside a fragment you cannot see the shape of what you were cut out of. Reasoning cannot recover information that was never delivered. A careful reader of a fragment reaches the fragment's conclusion.
Which is why the fix cannot live in the writing. Writing better pages produces better pages that get read in pieces.
The door
Here is what worked in my house. I am offering it as a shape rather than the shape, and at the end I will name the property that actually matters, so you can build a different version and still be fine.
I put a door in the root of the repository. One page, in one known place, listing the documents allowed to carry authority. If a document is not on that page it has no authority. Not less. None. A file that claims to outrank other files, in its own body, and is not on the door, is not a competing claim. It is history that happens to be legible.
The door is rude on purpose. A closed list is the only thing that beats a sentence claiming to win, because that sentence lives inside a document and the door sits outside all of them. You cannot argue your way on from inside a file. Somebody has to put you there.
Two kinds of document get on, and only two.
The first is law, meaning an append-only ledger. You add at the bottom and never rewrite an old entry. Each ruling gets a number and a date, and a ruling that changes an older one goes underneath, noting what it supersedes and leaving the old one in place, wrong, permanently. Law is trustworthy not because it is always right but because it is honest about its shape: a chronological record, highest number on a topic wins. A fragment of a ledger arrives with its number, its date and its ruling all inside the window, so it survives being read in pieces. Age is not a defect in a ledger. Age is the index.
The second is re-measurable, meaning the page carries a runnable command that proves its own claims. Not described. Printed, so a doubting reader can check in seconds. A page like that is allowed to be wrong, because it carries its own refutation procedure. The claim is not really that production holds four payment keys. It is: run this and see. One warning, because I got this wrong the first time: a proof is only worth its date, and the date has to be fresh against how fast its subject moves. The most confidently wrong page in my corpus printed its command and its output under a date that was honest when written.
Everything else is a snapshot: a picture of how the world was on a particular day. Architecture overviews, environment matrices, plans, phase notes. You cannot eliminate them and should not try. A snapshot gets a date, admits in its own voice that it is a photograph, and never goes on the door. One claiming to settle disputes is the most dangerous artifact in an agent system.
Now the objection, which is a good one. If agents write a hundred documents a day, a hand-maintained list cannot keep up.
It cannot, and it does not have to, because the door was never an index of the corpus. You are not cataloguing what exists. You are naming the handful of files permitted to say stop, and almost nothing needs the power to halt a human. A day's thirty documents are records, notes, plans and evidence, all useful, and not one needs the authority to stop a deploy.
So the flood is the reason the door is short rather than the reason it fails. My corpus could reach ten thousand documents without the door growing a line, because the number of things allowed to halt me scales with how many load-bearing systems I run, which is a low number and moves slowly. A dozen files is a job a person can do closely. Two thousand never will be.
Which settles where the person belongs. The human admits, and the machine audits. Neither half survives alone: a list nobody rechecks becomes the rotten page everybody trusts because it is on the list, and admission by machine alone is an index that grows with the corpus. Both entry rules are mechanical enough for a script to recheck on a schedule, without anybody remembering to.
Everything not on the door is silent by default. A document your agent writes at two o'clock has no authority at two oh one. It might be excellent, and more accurate than anything on the list. It still has none, because authority is not a property of accuracy, it is a property of admission.
Now the property, so you can build your own version. Authority has to be conferred from outside the artifact, so no artifact can confer it on itself, and it has to be retained by measurement rather than by having once been granted. A page in the root of a repository is how I arranged that. If your world is a wiki, a rules folder or a memory server, arrange it differently. A different shape that preserves the property is fine. What is not fine is authority a document assigns to itself, and authority nothing ever rechecks.
What it feels like when it works
The change is not that the documents become correct. The change is that alarms mean something again.
When my agent stops me now, the stop is information. I read it, I feel the appropriate amount of attention, and I act on it. The account it draws from, my willingness to take a warning seriously at eleven at night when I am tired, stopped being overdrawn, and that account turns out to have been the actual asset all along.
The second thing is less obvious and worth more. I move faster, because I can trust the stops. A gate you believe is a gate you can build on, so you stop hedging everything, stop keeping a private second opinion about your own system, and stop routing around your own safety rules because they have cried wolf too often. Speed came from trust, not from removing friction.
And the triage takes under a minute, which is the part you can use tomorrow morning whether or not you ever build a door.
A real gate cites one of two things. Either a measurement with a date, meaning a command was run, here is what came back, here is when. Or a ruling with a number you can find in a ledger. Nothing else counts as a citation. Not a page title. Not a heading announcing itself as the source of truth. Not the phrase per the documentation.
Anything else is citing a page that still thinks it is in charge.
So the question is not what it says, and it is not how sure you are. It is where the authority came from, and whether anybody has checked that source against the running world. If the source is trustworthy, measure anyway, because a check you never run is a check you do not have. If it is a ghost, proceed, and then move the ghost out of the path in the same session, while you still know where it is and why. Otherwise it stops you again in three weeks and you pay the same afternoon twice.
What it cost me
I am not going to tell you what this will cost you, because I do not know your setup and a hypothetical cost is fear dressed for work. I will tell you what it cost me instead, which I actually paid.
A month spent avoiding half an hour of work. Not a dramatic month. A month of ordinary days in which something more urgent came up, and the more urgent thing was always real, and the test never got run.
A product that would not ship, because every route to the last test ended in a stop I could not adjudicate. A project I stopped walking toward, and did not notice I had stopped walking toward, because it never presented as avoidance. It presented as priorities.
It cost me a month of a working product and a stretch of my life I did not enjoy.
And then the cleanup: two days, a framework built twice, and an audit on my own repair work that came back saying the claim was false, which was humiliating in the specific way that finding your own name on a bad decision is humiliating.
The line I would want somebody to have told me is this one. My repositories came back faster than my trust did. Cleaning the corpus was two days of work. Believing an alarm again, at eleven at night, without a reflexive flinch of here we go, took considerably longer, and I could not shorten it by knowing the cleanup had been done. Trust is rebuilt by alarms turning out to be real, one at a time, and there is no way to do that quickly.
What to do tomorrow
Put a door in the root of one repository. Not all of them. One. A single page listing the documents that are allowed to stop you.
Then be ruthless about what goes on it, because the ruthlessness is where the value is. For each candidate, ask whether it should be allowed to halt a deploy. The answer is almost always no, and every no is an afternoon you do not spend later.
There is a free packet if you want a running start. It holds the technical version of this piece written for agents to read, a door template, a ledger template, a template for a document that proves its own claims, a prompt that has an agent propose a door for your repository without changing anything, and a script that checks whether the documents on your door pass the admission test. It is on GitHub, it is free, and there is no email to give.
Now the limits, because I would rather you trusted the parts that hold. I do not know how far this generalizes past systems shaped like mine, meaning agents with a durable role, a body of loaded truth, and memory across sessions. A single-shot automation with a fixed prompt has no room for the failure, and where the boundary sits I could not draw for you. I do not know that the door survives an organization either. Mine is held by one person and works because that person reads it closely. I can picture a company where the list grows too long to read and too political to prune, and becomes the thing it was built to prevent.
Knowing about this does not make you immune. I named the failure, built the detector twice, wrote the framework, hold the door personally, and still found drift inside the memory system I built specifically to prevent it, while I was writing the argument against it. Awareness buys one advantage: you will recognize it faster when it happens to you.
Your corpus has sentences in it right now that were true when somebody wrote them and are not true anymore. So does mine. One of them is going to stop you this week, faithfully, with a citation, in the register you reserve for real danger.
Do not argue with it. Do not obey it either. Ask where the authority came from, and whether anybody has checked it against the running world. If the answer is nobody, what stopped you is exhaust, and the thing to do is move it out of the path so it cannot stop you a second time.