The work behind the model
Every AI system that recognises a face, transcribes a voice or answers a question in fluent English was trained on examples that people prepared by hand. Someone drew the box around the cyclist. Someone decided which of two chatbot answers was more helpful. Someone listened to four seconds of muffled audio and typed out what was said.
That work is called data annotation, and it is one of the few genuinely entry-level remote jobs attached to the AI industry. You do not need a machine learning degree. For most tasks you need careful attention, patience, and a willingness to read instructions properly and then follow them even when you disagree with them.
This lesson is about what the job actually is, day to day, so you can decide whether you want it before you spend a fortnight applying.
In plain English
- Annotation:
- Adding labels, ratings or corrections to raw data so a model can learn from it. Also called data labelling.
- Guidelines:
- The written rules for a project, often dozens of pages, defining exactly how each label should be applied.
- Quality score:
- A number the platform keeps on how often your work matches the reviewer's answer. It usually decides whether you keep getting tasks.
- Queue:
- The pool of available tasks. When it is empty there is no paid work, however much time you had set aside.
- Audit:
- A sample of your completed tasks re-checked by a reviewer, which feeds your quality score.
The main kinds of task
Image and video labelling. Drawing boxes or outlines around objects and tagging what they are: road signs, products, tumours, defects on a production line. Repetitive and precise.
Rating and comparing AI responses. You are shown a question and two model answers, and you judge which is better against a written standard covering helpfulness, accuracy and safety. Sometimes you write the better answer yourself.
Transcription and audio. Typing what is said in a recording, tagging speakers or accents, or recording your own voice reading prompts in your language or dialect.
Search and fact checking. Rating whether a search result actually answers a query, or checking whether a claim a model made is true and finding the source that proves it.
Content classification and safety. Deciding whether a post, image or message breaks a specific written rule.
Specialist review. Checking model output in a field you already know: code, medicine, law, accounting, a less widely spoken language.
What a shift is actually like
You log into a platform, open the queue and work through tasks. Each one has a target time. Before you begin you read a guidelines document that may run to forty pages, and you will go back to it constantly, because the difficult part is never the obvious cases. It is the edge cases. Is a car reflected in a shop window a car? The guidelines have an answer, and your job is to give their answer rather than yours.
The guidelines are the job. People who struggle with annotation work are usually not careless. They answered sensibly instead of answering the way the document said to.
Your work gets audited. A reviewer re-checks a sample and your quality score moves. Below a threshold you are retrained, taken off the project, or you simply stop receiving tasks. You are not always told which.
RULE 4.2 PARTIALLY VISIBLE OBJECTS
Label an object if 25 percent or more of it is visible.
If less than 25 percent is visible, do not label it.
Reflections and images on screens are NOT real objects. Do not label them.
Objects inside a printed advertisement ARE labelled, using the tag ADVERT.
Most common errors in last week's audit:
- Labelling reflections in shop windows (rule 4.2)
- Boxes drawn loosely instead of tight to the object edge
Two other realities worth knowing early. Work is not steady: queues empty without warning, so a busy week can be followed by one with almost nothing in it. And the pace is timed, so you are measured on speed and accuracy at once.
The part that is rarely in the advert
Some annotation work, particularly anything near content moderation or model safety, means looking at material chosen precisely because it is harmful. That can include graphic violence, sexual content including material involving children, self-harm, and extremist propaganda. Companies do hire for this, and the job ad does not always say so plainly.
You are allowed to refuse this work, and you should decide before you accept rather than on day three. Repeated exposure to distressing material causes real psychological harm and is a recognised occupational risk in this industry. Ask directly what kind of content a project involves, ask what support is provided, and treat a vague answer as an answer.
If you do take safety-related work, ask about breaks, blurring or greyscale viewing tools, caps on daily exposure, and access to counselling. Some employers provide these. Many do not.
Checkpoint
Annotation means following someone else's written rules exactly, at speed, while being scored on accuracy. Always ask what content a project involves before accepting it.
- What exactly will I be looking at or listening to?
- Does this project include violent, sexual or self-harm content?
- How long are the guidelines, and am I paid for reading them?
- Is pay per task or per hour, and is training time paid?
- How is my quality score calculated, and can I see it?
- What happens to my pay if a task is rejected?
- How many hours of work are realistically available per week?
Set a timer for 20 minutes. Sort a folder of your own photos into
three groups using a rule you write down FIRST, for example:
A = at least one person clearly visible
B = no people, outdoors
C = no people, indoors
Extra rule: a reflection of a person does not count as a person.
At the end, answer honestly:
- Did you keep applying your written rule, or drift into common sense?
- Was 20 minutes of this fine, or already tiring?
- Could you hold the same accuracy for three hours?
That last question is the real one. Plenty of people do this work well and find it steady and manageable. Plenty of others discover after two weeks that the repetition costs them more than the pay is worth, and there is nothing wrong with being in the second group. Better to learn it in twenty minutes than in a month.
๐ Quiz
Question 1 of 4The guidelines say reflections in windows are not real objects, but you can plainly see a car reflected in a shop front. What do you label?