The short answer
Most people using AI at work have never checked whether it is actually saving them time. They feel faster, which is not the same thing, and the honest test takes about two weeks: pick one recurring task, remember roughly how long it used to take you, do it with AI for a fortnight while noting the total time (including all the fixing), and compare finished work to finished work. Where AI wins, keep it. Where it loses, drop it without guilt.
That is the whole method. No spreadsheet, no dashboard, no time-tracking app. The rest of this post is about why the feeling of speed misleads almost everyone, what the two classic failure patterns look like, and how to run the test without turning it into a project of its own.
Why using AI always feels productive
AI tools are, among other things, extremely good at feeling fast. You type a request and text pours onto the screen immediately. Something is always happening. Compare that to staring at a blank document for ten minutes, which feels like nothing is happening even when your brain is doing the most important part of the work.
The trouble is that the visible part is not the whole task. The full loop includes writing the prompt, reading the output, deciding it is not quite right, prompting again, reading again, then fixing the tone, checking the claims, and rewriting the bits that sound nothing like you. Every one of those steps is easy to forget afterwards, because each one felt quick. The draft arrived in thirty seconds. The finished, sendable piece may have taken longer than it ever did before, and you would genuinely not know.
In plain English
- Finished:
- Approved and sent, or good enough that you would put your name on it. Not 'first draft produced'. All honest comparisons use this definition.
- Review time:
- The minutes you spend reading, checking and correcting AI output. Real work, and the part everyone forgets to count.
- Baseline:
- Roughly how long the task took you before AI. Without some idea of this, you have nothing to compare against.
The two-week test
Here is the whole procedure, and it deliberately fits inside a normal working life.
Pick one task. Something you do at least weekly: the status update, the client email, the meeting summary, the monthly report. One task, not your whole job. If you try to measure everything you will measure nothing.
Recall your baseline. Think about the last few times you did it the old way and settle on an honest rough figure. Twenty minutes? An hour? Precision does not matter here. You are not writing a business case, you are answering a yes or no question, and "about forty minutes, usually" is plenty.
Do it with AI for two weeks. Each time, jot down the total time from starting to genuinely finished. A note on your phone is fine. The only rule is that the clock includes everything: prompting, re-prompting, reading, checking and fixing.
Compare finished to finished. Not the old finished version against the new first draft. The thing you would actually send, both times.
Then act on what you find. This is the step people skip. If AI made your meeting summaries faster and your client emails slower, the answer is not "AI is good" or "AI is bad". The answer is: keep it for summaries, write the emails yourself, and feel no guilt about either.
Dropping AI for a task it loses at is not falling behind. It is the entire point of testing. The people who get real value from these tools are the ones who use them selectively, and you can only be selective if you have checked.
The two ways the test comes back negative
When AI loses the comparison, it almost always loses in one of two ways, and both are worth recognising by name.
Fixing time swallows the drafting time saved. The draft took two minutes instead of twenty, but the checking and correcting took twenty-five. This happens most on work where the details matter and the details are yours: specific figures, client history, anything the tool cannot know and will confidently guess at. The saving on typing is real, and it is smaller than the new cost of verification.
Scope creep. The AI version is longer, has more sections, includes an executive summary nobody requested, and took ninety minutes of shaping instead of the thirty minutes the plain version used to take. Nothing about it is wrong, exactly. It is just more than anyone needed. Because generating extra material is nearly free, the temptation to include it is constant, and the reader pays for it in reading time even when you do not pay in writing time.
Both patterns share a root: the tool changed the task instead of just speeding it up. Watch for that. If your weekly update was three paragraphs before AI and is two pages now, you have not become more productive. You have become more prolific, which your colleagues may already have noticed.
Where AI tends to win and lose
Your own test outranks any general claim, including these. But as a starting point for choosing which task to test first, this is roughly how it falls.
| Usually saves time | Often loses time |
|---|---|
| First drafts of routine writing | Anything where the facts need checking |
| Summarising long documents and meeting notes | Arithmetic and figures |
| Reformatting: prose to bullets, notes to tables | Short replies faster to write yourself |
| Rewriting for tone or a different audience | Judgement calls dressed up as writing tasks |
| Getting unstuck on a blank page | Work where the details live only in your head |
The right-hand column is not "things AI cannot do". It is things where the checking burden tends to eat the saving, which is a different and more useful claim.
Do not trust anyone's numbers, including ours
You will meet plenty of confident statistics about how much time AI saves the average worker. Treat them all the same way: politely, and from a distance. They describe other people, other tasks, other definitions of finished, and they are usually measured in ways that flatter the answer.
The wonderful thing about the two-week test is that it makes every one of those numbers irrelevant to you. "Did it help" is a question you can settle for yourself, about your actual work, in a fortnight, for free. That beats any figure anyone quotes at you, in either direction.
If you want a light structure for the end of each test week, this works well:
I have been testing AI on one recurring task this week. Here are my notes:
Task: [what it is]
Rough time the old way: [X minutes]
Times this week, start to genuinely finished: [list them]
What I had to fix each time: [notes]
Tell me plainly: does this look like a real saving, once fixing time
is counted? Was the output longer or fancier than the old version
needed to be? Do not encourage me. Just read the numbers.
Where to go next
If the test tells you AI earns its place, the next step is pointing it at more of your week properly, which is what AI Productivity at Work covers. And if you manage people and need the team-level version of this question, with baselines and business cases and what to say to finance, that is AI for Managers and Team Leads. Both are free, and both take the same position this post does: measure your own, and believe what you measure.