LearnAI home

Making It Work ยท Lesson 3

How to Choose Your Team's First AI Use Case

Pick the boring one that works, not the impressive one that fails.

Your first use case is a governance decision

It does not feel like one. It feels like a productivity choice, so people pick the task with the biggest apparent saving, or the one that will look best when somebody senior asks how adoption is going.

But the first use case sets precedent. Whatever you choose becomes the working example of what AI is for here, how much checking is normal, and how much trust the output gets. Choose something that goes straight to customers and you have taught your team that unreviewed AI output can leave the building. That lesson will outlast the pilot.

So the right question is not "where would AI save the most time". It is "where can we be wrong cheaply while we learn what wrong looks like".

In plain English

Blast radius:
How far the damage travels if an output is wrong and nobody catches it.
Reversible:
A mistake you can quietly correct, as opposed to one already sitting in a customer's inbox.
Baseline:
How good or fast the work was before AI touched it. Without one you cannot tell whether anything improved.
Draft tolerant:
Work where a rough first version is genuinely useful because a person edits it next.

The shape of a good first choice

The candidate you want is boring on purpose. Four features, and you want all four.

High volume. It happens weekly or daily. Volume is what turns small savings into real ones, and it is also what lets you see patterns in the errors. Do something twice and you learn nothing about how it fails.

Low judgment. The work is assembly, formatting, summarising or first drafting. If the value of the task is somebody deciding a hard thing, AI cannot touch the part that matters.

Draft tolerant. A rough version is useful because a human edits it next. If the first version has to be right, you have moved effort rather than saved it.

Small blast radius. If it is wrong and slips through, the damage is internal, visible quickly, and fixable without a phone call to a client.

Add the four together and you get the kind of task nobody puts in a slide deck. Meeting notes into actions. Long documents into summaries. Rough notes into a consistent format. A first draft of a recurring internal update. A set of similar records tidied into the same structure.

If describing the use case out loud makes it sound unimpressive, that is a good sign. Impressive usually means visible, and visible means the errors are visible too.

Three bad first choices

These are the ones teams reach for, and each fails for a different reason.

Customer facing work without a review step. Chatbots answering live, replies sent automatically, anything reaching an outsider without a named person reading it first. The blast radius is external and the correction is public. If you want to use AI on customer replies, and many teams should, start with drafts a human sends, not messages a system sends.

Anything with legal or contractual consequences. Contract wording, formal notices, regulatory submissions, statements about what your company is obliged to do. Being confidently wrong here is not embarrassing, it is expensive, and it puts your legal function in the position of finding out afterwards.

Anything nobody measured before. This one is quieter and it kills more pilots. If you cannot say how long the task took before, or how often it was wrong before, you will have no way of showing whether AI helped. Six weeks later somebody senior asks for the result and you have anecdotes. Pick something with a number attached, even a rough one.

โŒ Weak prompt

Prompt

What is a good first AI use case for my operations team?

Output

Consider automating customer support responses, generating reports, forecasting demand, and streamlining data entry across your workflows.

Four categories, no verdict, and the first suggestion is the highest blast radius option available. You cannot start any of these on Monday.

โœ… Good prompt

Prompt

Here are five recurring tasks my operations team does each week, with roughly how long each takes and who reviews the output today. Score each one against four tests: high volume, low judgment, tolerant of a rough draft, and small blast radius if an error slips through. Give a verdict of start here, later, or avoid, with one sentence of reasoning. Do not suggest tasks I did not list.

Output

Weekly supplier summary: start here. Weekly, mostly assembly, always read internally before it goes anywhere. Customer refund decisions: avoid. Judgment is the task, and an error reaches the customer directly.

Your real tasks, four explicit tests, and a decision at the end. This is a shortlist, not a brochure.

Prompt you can copy: score your candidate tasks

Here are the recurring tasks my team does, with rough weekly hours and who currently checks the output: [task 1, hours, who checks it] [task 2, hours, who checks it] [task 3, hours, who checks it] [task 4, hours, who checks it]

Score each against four tests:

  1. High volume (weekly or daily)
  2. Low judgment (assembly, formatting, summarising, first drafting)
  3. Draft tolerant (a rough version is useful because a person edits it)
  4. Small blast radius (if it is wrong and missed, the damage is internal and quickly fixable)

Give each a verdict of start here, later, or avoid, with one sentence of reasoning. Do not invent tasks I did not list.

Checkpoint

Your first use case sets the precedent for how much checking is normal. Pick boring: high volume, low judgment, draft tolerant, small blast radius. Avoid customer facing work without review, anything with legal consequences, and anything nobody measured before.

Write down the baseline before you start

Five minutes of work that almost everyone skips. Before the first prompt, write one line: this task takes roughly this long, happens this often, and goes wrong about this much. It does not need to be precise. It needs to exist, because a rough number written down beforehand beats a precise number reconstructed afterwards by someone with an opinion.

Prompt you can copy: define the baseline and the review date

The task we are starting with is: [describe the task in one sentence, including who does it and how often]

Help me write a short baseline note, no more than 150 words, covering:

  • How we will measure time spent, before and after
  • What counts as an error for this task, in concrete terms
  • Who reviews the output before it is treated as finished
  • What we would need to see to keep going, and what would make us stop

Then suggest a sensible date to review it. Assume we want a real answer, not a positive one.

Never let a vendor demo pick your first use case. Demos are built to look impressive, which is the exact opposite of the selection criteria that matter. Start from your own recurring work, then look for a tool.

Name one person, not a committee

Give the first use case an owner. Not a working group, one person who runs it, keeps the baseline note, and has authority to stop it. Committees are excellent at starting pilots and terrible at ending them, and knowing how to stop cleanly is what makes a team willing to try the next thing.

Prompt you can copy: argue against your own choice

We have chosen this as our team's first AI use case: [task, who does it, how often, who reviews the output]

Argue the case against it. Specifically:

  • What happens if an output is wrong and nobody notices for two weeks?
  • Who would have to answer for that, by role?
  • Does any client or personal information have to be entered to do it?
  • Could we tell in six weeks whether this actually helped?

Be blunt. If this is a poor first choice, say so plainly and say why.

Asking for the case against your own decision is the cheapest review available. The objections you cannot answer are the ones worth acting on.

๐Ÿ“ Quiz

Question 1 of 4

Why is the first use case treated here as a governance decision rather than a productivity one?

Found this useful? Pass it on.