An AI assistant for long documents, analysis and code, with Claude Code and Cowork on paid plans.
AI evaluator and QA tester
New roleChecks how an assistant or agent answers before launch and after every change: builds test questions, hunts for errors, made-up answers and ways to trick the system. Without that kind of check, customers are the ones who report the problems.
Where this job sits on the map
In the same group: 14 of 100 professions.
What changes
You can't check a model's answers once and be done: when the model or its instructions change, the whole set of examples is rerun.
Attack testing is now part of the job: attempts to make the assistant break its rules or reveal data (red teaming).
Part of the grading is done by another model, but the grading rules and the borderline cases stay with a person.
AI drafts, the person decides
How the work splits here: AI prepares a draft, the person checks it and makes the call.
What AI does
- Comes up with many test questions, including trick ones
- Grades answers against set criteria
- Compares two versions of an assistant on the same set of questions
- Groups errors by type and writes a summary
What stays with the person
- Deciding what counts as a right answer in this field
- Test cases drawn from your customers' real situations
- The verdict: "ready to launch" or "too early"
- Describing the risks in plain language for management
What to learn first
6 lessons from the course, in order. Start with the first one.
- How an LLM works inside, explained without mathStarting from zeroHow a language model works inside: tokens, context, temperature and hallucinations, explained without math.Start here
- How to write a good promptUserHow to give AI a task: the five parts of a good prompt, Plan Mode, and how to refine an answer instead of starting over.
- AI ethics and safety: hallucinations, attacks, biasUserHallucinations, prompt injection, bias, privacy and copyright: how not to trust AI blindly.
- MLOps for indie builders: monitoring, drift and retraining without a DevOps teamEngineerMLOps without a DevOps team: call tracing, prompt regression tests, cost control and prompt versioning.
- Evals: skills that improve themselvesEngineerEvals for skills: tests, pass rate, checking that the skill triggers and a data-driven improvement loop.
- Prompt injection defense, the top security threat of 2026EngineerPrompt injection: six attack types, eight layers of defense, red team tests and an incident response plan.
Which tools to use
- Freemium
OpenAI's mainstream AI assistant: chat, voice, images, the Work agent and Codex.
FreemiumGoogle's assistant, built into Gmail, Docs and Sheets, with Deep Research, Gemini Live and Canvas.
FreemiumGoogle's free studio for Gemini: Build mode creates an app from a description or sketch.
FreemiumGemini in Google Sheets: tables, formulas and data analysis from a description.
Paid
Ready-to-use materials
- Check AI output before you publishChecklist
Facts, meaning, risks and voice: what to check before customers see text written by AI.
Free - How to check an AI answerChecklist
Assess the risk, spot the signs of a made-up answer and run a 5-minute check before you act on what AI says.
Free Go through an AI answer claim by claim: what's right, what's doubtful, where it's wrong and what to check at the source.
Free- Data security when working with AIChecklist
What not to paste into a chat, which settings to check and how to limit AI agents.
Free
How to earn with AI
We don't promise income: results depend on your niche, your market and your work.
- Job
If you're a tester or work in customer support, take on checking your company's AI assistant: question sets, bug reports, repeat runs.
- Service
Testing a chatbot or assistant before launch: a question set, a report on errors and risks, and a recheck after the fixes.
- Product
A ready-made set of test questions and criteria for assistants in one industry, for example online stores.
Similar professions
Checked: October 2026