Getting Paid to Break AI: Red-Teaming and Safety Evaluation Work
2026-09-22 · Expert Match AI team
Every frontier lab runs a red team - people whose job is to make the model do things it should not, so the lab can fix it before a user finds the same hole. For a long time that was an in-house function with a security-research flavour. Over the last year it has become a category of contract work, and it is now showing up on the same platforms that list attorney and physician queues. It just does not show up under the name you would search for.
We checked all 316 roles on our board. Not one has "red team" in the title. Around a dozen are red-teaming or safety evaluation work by any reasonable reading of the description. This post is how to find them, what they pay, and whether you would want to.
What the work actually is
Safety evaluation splits into three tasks, and they suit different people.
Adversarial prompting. Writing the inputs that try to get a model to produce harmful output - the jailbreak, the roleplay framing, the multi-turn setup that gets around a refusal. The lab wants a catalogue of what works so it can train against it. This is the part people picture, and it is the smallest part of the paid market: the labs' own teams and a few specialist firms do most of it, and the platform work is usually scoped to a narrow domain rather than open-ended attack.
Harm evaluation. Reading model outputs and judging whether they cross a line: is this medical answer dangerous, does this response to a distressed teenager make things worse, does this legal answer expose someone to real liability. This is graded work with a rubric, closer to RLHF rating than to hacking, and it is where most of the listed roles sit. It needs domain judgment more than technical skill.
Policy and rubric authoring. Writing the standards the graders use: what counts as harmful in this domain, where the line is, what a good refusal looks like versus an unhelpful one. This is the senior tier, and it is the work the credentialed specialist listings are really recruiting for.
What is on the board right now
The live hook is micro1, which over the first two weeks of September posted a cluster of clinical safety specialist roles that are unmistakably harm-evaluation and policy work:
| Role | Listed rate | Posted |
|---|---|---|
| Adolescent/Child Safeguarding & Exploitation Specialist | $45–95/hr | Sep 1 |
| Eating Disorder Specialist (Adolescent) | $45–95/hr | Sep 1 |
| Substance Use – Adolescent Addiction Specialist | $45–95/hr | Sep 1 |
| Adolescent Suicide & Self-Harm Specialist | $45–95/hr | Sep 2 |
| Military/Veteran Mental Health Experts | $100–200/hr | Aug 25 |
| Developmental Psychology (PhD) | $100–200/hr | Aug 25 |
| Bilingual Mental Health Expert | $100–200/hr | Sep 14 |
| Child & Youth Abuse Trauma Specialist | $100–200/hr | Sep 14 |
Read those together and the shape of the commission is obvious: a lab is building or auditing how its model responds to minors and vulnerable users in crisis, and it wants licensed clinicians to grade the responses and write the standards. The $45–95 tier is evaluation; the $100–200 tier is the senior rubric and policy layer, priced like the rest of the credentialed expert market.
Elsewhere: Braintrust's RLHF Healthcare Expert ($100–180) is harm evaluation in a clinical domain by another name; micro1's Data Privacy Analyst ($31–60) is the compliance-flavoured edge of the same work; and Alignerr's healthcare call reviewer roles ($15–20) are the annotation-tier version - reviewing recorded calls for a healthcare AI against a QA standard.
Sort the jobs board by pay and search the titles for safeguarding, mental health, trauma, privacy and RLHF and you will find the current set; the platforms do not tag it.
Who the labs want
Not, on the evidence of the board, hackers. Every safety listing we can see asks for a clinical or professional credential first: licensed therapist, psychologist, social worker, addiction counsellor, PhD in a relevant field. The lab's problem is not generating attacks - it can do that itself - but knowing which outputs are actually harmful to a real person in a real situation, and that is domain judgment.
The exceptions are the technical adversarial roles, which are rarer on the open platforms and tend to appear as "LLM evaluation" or "agent evaluation" engineering listings with a safety component in the description, rather than as safety roles. If you are a security researcher, that is where to look, and the engineering evaluation post covers the queue.
The temperament question
This deserves a plain statement. Harm evaluation means reading, for hours, model outputs about the worst things that happen to people - and, in the adolescent-safety cluster above, specifically about children. The labs are recruiting clinicians for it partly because clinicians already do this professionally and have the training and the supports to do it sustainably. If you do not, the rate is not a reason to start. Several platforms include a content warning in the onboarding for these queues; take it at face value.
For the people who do this work already, the case for the platform version is decent: it pays at or above clinical contract rates, it is remote and asynchronous, and the work - deciding what a model should say to a fourteen-year-old in crisis - has more downstream reach than any individual caseload. That is the honest pitch, and it is the one the listings are making.
Getting in
The gate is the credential, then the platform's standard screen. micro1 routes these applicants through its AI interview like everything else, so the credential string and the domain specifics need to be said out loud. Because these are batch-shaped commissions - eight roles posted across two weeks for what is plainly one project - they will close when the pool fills. If you qualify, this is not a queue to think about for a month.
The salary report tracks the healthcare and science categories these roles fall into, and uploading a resume on the homepage will tell you which of them your licence actually matches.
Published by the Expert Match AI team. Role titles, rates and posting dates are from our tracker as of September 15, 2026 (316 roles tracked, 10 platforms) and are listed rates, not accepted offers; our reading of what each role involves is based on the published description and may differ from the platform's internal scoping. Some outbound application links on this site carry disclosed referral codes; rankings and recommendations are never influenced by referral payout.