Why most review processes fail quietly
A team introduces AI drafting. Someone sensibly says everything should be reviewed. A manager agrees to check the output. For two weeks the checking is genuine. By week four it is a scroll and a nod. By week eight nobody remembers it was supposed to be happening.
The process did not collapse because people are lazy. It collapsed because "review everything" gives no instruction about what to look at. Faced with a page of fluent, well organised, confident prose, a reader with no specific target reads for tone and structure, finds both excellent, and moves on. Fluency is exactly what AI is best at, so the surface always passes.
Reviewing everything and looking for nothing catches nothing.
In plain English
- Targeted review:
- Checking a defined list of high risk elements rather than reading the whole thing for a general impression.
- Author check:
- The first verification, done by the person who produced the draft, before it reaches anyone else.
- Spot check:
- A manager reading a small sample deliberately, looking for patterns rather than individual mistakes.
- Systemic error:
- A mistake that repeats because of how the work is set up, not because one person slipped.
Check four things, properly
Replace "review everything" with a list short enough to hold in your head.
Numbers. Every figure, percentage, amount and count, verified against the source it came from. Not sense checked, verified. A wrong number that looks reasonable is the single most common way AI assisted work goes wrong, because reasonableness is what the model optimises for.
Names. People, companies, products, job titles. Models confuse similar names, promote people, and merge two organisations into one plausible hybrid.
Commitments. Anything the draft promises on your behalf. Dates, deadlines, deliverables, next steps, "we will". A draft that invents a Friday deadline creates a real obligation the moment it is sent.
Claims. Statements about what your company does, offers, guarantees or has done. These are the ones that cause problems long after the document is forgotten.
Four categories. Ninety seconds on a short document. And crucially, each one is a specific action with a right answer, which is why it survives contact with a busy week in a way that "read it carefully" never does.
Tone and structure need no review process. Those are the things AI does well and the things a human notices automatically. Spend the attention where the failure actually lives.
Push the first check to the author
The instinct is to make the manager the checkpoint. It is the wrong design for three reasons.
The author has the source material open. The manager does not, so the manager can only assess plausibility, which is precisely the quality the model already maximised.
The manager becomes a bottleneck, and bottlenecks are routed around. Work gets sent late on a Friday. Work gets marked urgent. Eventually work stops arriving at all, and you find out months later.
Most importantly, making the manager the check tells the author their draft is somebody else's problem. That is the exact opposite of "the person who sends it owns it", and your team will follow the process you built rather than the principle you announced.
So: the author runs the four checks, every time, before anyone else sees it. The manager does something different.
โ Weak prompt
Prompt
Review this draft and tell me if it is good.
Output
This is a well structured and professional draft. The tone is appropriate and the key points come across clearly. A few minor wording suggestions below.
It graded the writing, which was never at risk. Every figure and every claim went unexamined.
โ Good prompt
Prompt
Do not comment on tone, structure or wording. Extract from the draft below, as four separate lists: every number and what it refers to, every person or organisation name, every commitment or deadline it makes on my behalf, and every claim about what my company does or offers. If anything appears in the draft that is not in the source notes I have pasted underneath, flag it as UNSUPPORTED.
Output
Numbers: 14 percent, referring to renewal rate. UNSUPPORTED, not present in source notes. Commitments: report delivered by Friday 12th. UNSUPPORTED.
It produced a verification list, not an opinion. The unsupported flags are the two things that would have caused a problem.
Below is a draft, and underneath it the source notes it was written from.
Do NOT comment on tone, style or structure.
Produce four lists:
- NUMBERS: every figure, percentage, amount or count, and what it refers to
- NAMES: every person, company, product or job title
- COMMITMENTS: every date, deadline or promise made on my behalf
- CLAIMS: every statement about what my company does, offers or has done
Mark anything that does not appear in the source notes as UNSUPPORTED.
DRAFT:
[paste draft]
SOURCE NOTES:
[paste notes]
Checkpoint
Reviewing everything and looking for nothing catches nothing. Check four things properly: numbers, names, commitments, claims. The first check belongs to the author, who has the source open. Managers sample for patterns instead.
What the manager actually does
Not everything. A deliberate sample, read slowly, with one question in mind: is there a pattern here?
Take five pieces of output a fortnight. Read them properly, source material open. You are not trying to catch every error, which is impossible and would make you the bottleneck you were avoiding. You are trying to find out whether the same kind of error keeps appearing.
That distinction changes what you do next.
One off or pattern
A one off is a person who was rushing. The response is a quiet word and a reminder of the four checks. No process change, no all staff email. Reacting to a single slip with a new procedure is how organisations accumulate rules nobody follows.
A pattern is a design fault, and you must resist the urge to fix it by asking people to try harder. If figures are wrong repeatedly, the source data is probably not in front of the author when they draft. If commitments keep appearing that nobody made, the prompt template needs an explicit instruction not to invent them. If a whole category of work keeps failing, the use case may be wrong.
Three of the same error is a system talking to you.
Here are the errors we found in the last [ten] pieces of AI assisted work,
with a short note on each:
[error 1]
[error 2]
[error 3]
Group them by underlying cause, not by symptom. For each group, tell me:
- Whether this looks like individual slips or a repeating design fault
- The single most likely reason it keeps happening
- One change to the process, prompt or source material that would fix it
Do not recommend telling people to be more careful.
If your review process has never rejected anything, it is not a review process. Something that passes everything is decoration, and it gives you confidence you have not earned.
Our first AI use case is:
[task, who does it, how often, who it goes to]
Write a review step for it in under 120 words, aimed at the person doing
the work. It must:
- Name exactly what they check, in specific terms for this task
- Say who does the checking and at what point
- Say what happens when something is found
- Fit in under two minutes per item
Plain language. No policy voice. Do not use the words diligence or robust.
๐ Quiz
Question 1 of 4Why does an instruction to review everything tend to catch nothing?