What the work actually is
"AI trainer" covers a family of tasks that share one skill: applying a written rubric consistently to messy material. Depending on the project, you might rate search results for relevance, or compare two model responses and justify a preference. You might write the reference answer a model should have given, label images or documents, or probe for unsafe outputs. The rubric is always long, the edge cases are the job, and your agreement rate with other raters is measured constantly.
Core tasks
- Rate responses for accuracy, tone, and policy compliance against a detailed guideline document.
- Write ideal reference answers for ambiguous prompts.
- Compare paired responses and explain the preference — the explanation matters as much as the pick.
- Red-team prompts for safety edge cases.
- Participate in calibration sessions where reviewers align on hard examples.
Who is actually hiring, and through whom
Most rating and labeling work routes through vendor platforms rather than the AI labs directly. The names are TELUS Digital, Appen, Outlier, DataAnnotation, and smaller shops like Workada. TELUS Digital absorbed Lionbridge's rater programs, so searches for "Lionbridge AI jobs" now land there. You contract with the vendor, and the end client is usually unnamed. That matters practically: your relationship, payment terms, and quality metrics all live with the vendor, and projects can end without notice when the client's need changes.
What real listings pay right now
Two dated data points come from our curated jobs page, both fetched July 2026 from Remotive with links to the source. TELUS Digital's US search-quality reviewer role pays $14/hr task-based, and Workada's data labeling runs $18–22/hr freelance. Treat these as the general-work floor. Domain-expert evaluation — medical, legal, or code review — pays multiples of that. But those projects screen hard for verifiable expertise, and the supply of generalist raters keeps generalist rates low.
How to get hired
Qualification exams test instruction-following and attention to detail more than brilliance. Read the guideline document twice before answering, because most failures are people skimming rules they think they already understand. Highlight domain expertise — healthcare, law, coding, languages — and note that many projects require a specific country of residence for locale-sensitive rating. That is why listings say "United States" or "Canada (French)" rather than "worldwide".
Watchouts
- Unpaid assessments are normal up to roughly an hour; multi-day unpaid "trials" are not. Walk away.
- Task-based pay quoted as an hourly figure assumes a task speed you may not hit for weeks.
- Read contracts for exclusivity clauses and data-handling rules before signing.
- Project pipelines dry up abruptly. Treat this as supplemental income until a vendor has kept you busy for months.
- Never pay anyone to "register" you for rating work — legitimate vendors do not charge applicants.