Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teamer — Adversarial Prompting & Evaluation

Join OpenTrain as a part-time AI Safety Red Teamer to design adversarial prompts and evaluate frontier models across high-risk domains; remote, 20+ hrs/week, paid $70–$84/hr. Requires 5+ years' relevant experience and strong written analytical skills.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $70–$84/hr

$70–$84/hr

Compensation

33 countries

Eligibility

Entry

Experience

Jul 17, 2026

Posted

Open to applicants in

United States Denmark Estonia Finland Ireland Latvia Lithuania Norway Sweden Austria Belgium France Germany Netherlands Switzerland United Kingdom Albania Bosnia & Herzegovina Croatia Greece Italy Malta Portugal Serbia Slovenia Spain Bulgaria Czechia Hungary Moldova Poland Romania Slovakia

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the leading platform for building careers in AI training and data labeling. We help people discover projects, build a unified portfolio of AI training work, and grow a durable freelance career teaching AI systems how to behave.

OpenTrain AI is the hiring and contracting organization for this role; you will join a distributed team that supports AI safety research and red-teaming work for frontier models.

Why AI training and red teaming matters

AI training (also called data labeling or human feedback work) is the human side of building artificial intelligence. Contributors design tests, label outputs, and evaluate model behavior so AI systems are safer and more reliable.

Red teaming is a high-impact part of that work: by probing weaknesses, uncovering jailbreaks, and documenting failures, you directly shape model alignment, robustness, and safety practices across the field.

The role — what an AI Safety Red Teamer does

You will stress-test frontier AI systems through adversarial prompts and structured evaluations. Work is text-focused and centers on difficult, high-risk, and gray-area topics such as misinformation, cyber threats, biosecurity, fraud, political content, and other sensitive domains.

This is a contract, part-time role (20+ hours/week) with flexible scheduling. The work is remote but limited to applicants based in the listed eligible countries.

  • Design adversarial prompts to probe model weaknesses and discover jailbreaks.
  • Evaluate AI outputs for unsafe behavior, hallucinations, policy failures, and risky content.
  • Document vulnerabilities and produce precise write-ups for benchmarking and red-teaming reports.
  • Collaborate with AI researchers to support alignment, robustness testing, and RLHF-related analysis.

Compensation, schedule, and employment type

This is a contractor, part-time position. Expect to work 20+ hours per week with flexible scheduling agreed upon with OpenTrain AI.

Pay is hourly, USD 70–84 per hour (typical rate shown: $84/hr; minimum $70/hr).

  • Employment types: Contractor, Part-time
  • Pay type: Hourly, USD 70–84/hr
  • Weekly time expectation: 20+ hours/week
  • Data type: Text; label types: RLHF, Evaluation Rating, Red Teaming

Required qualifications

Candidates must meet the following minimum requirements. OpenTrain will evaluate applications based on demonstrated experience and writing ability.

We will not invent or assume qualifications beyond what is listed below—please apply only if you meet these criteria.

  • Bachelor’s degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.
  • 5+ years of professional experience in AI safety, AI red teaming, trust & safety, cybersecurity, investigative journalism, life sciences, or similar fields.
  • Direct experience designing adversarial prompts or evaluating frontier AI systems and familiarity with jailbreak testing.
  • Practical familiarity with RLHF, SFT, AI alignment, or adversarial evaluation methods.
  • Strong analytical reasoning and excellent written communication; able to produce clear, precise vulnerability reports.
  • Ability to reason and write about sensitive domains such as cyber threats, biosecurity, misinformation, fraud, or political content.

Who should apply

This role is for experienced practitioners who enjoy adversarial thinking, have a sharp analytical mind, and can translate discoveries into well-written reports that researchers can act on.

Though the work is remote and flexible, applicants must be located in one of the eligible countries and able to commit 20+ hours per week.

  • Languages: English fluency required (writing and reading).
  • Eligible countries include US and a broad set of European countries (see application for full list).
  • Ideal backgrounds: AI safety researchers, security analysts, investigative journalists, life-science analysts, trust & safety experts.

How to apply and what to prepare

Apply through OpenTrain with a CV/resume and a short cover note describing relevant red-teaming or adversarial evaluation experience. Include concrete examples of past work (prompt designs, write-ups, published analyses, or evaluation tasks) if available.

Selected applicants will complete a skills assessment and sample red-team write-up as part of the evaluation process.

  • Provide examples that show prompt engineering, adversarial testing, or safety evaluation skills.
  • Be prepared for a short paid or unpaid skills check depending on assignment instructions (details provided after application).
  • OpenTrain will contact candidates who advance with next steps and scheduling for a reviewer or researcher interview.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Safety Red Team Specialist (English & Bengali)

Join OpenTrain to red team conversational AI: probe models with adversarial prompts, document reproducible attacks, and help improve safety. Part-time contractor role (20+ hrs/week), $20–22/hr, requires expert red teaming experience and fluency in English and Bengali.

Generative AI & RLHF
Text
Remote · Worldwide
English, Bangla
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026

Red-Teaming QA Lead (Remote US Contractor)

Join OpenTrain as a Red-Teaming QA Lead to oversee quality across safety and adversarial-evaluation projects; remote US contractor role, 20+ hrs/week, up to $100/hr. Lead reviewers, audit adversarial prompts and model outputs, and strengthen scalable QA for red-teaming work.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $100/hr

Posted Jul 9, 2026

AI Safety Red Teamer (English & Odia)

Join OpenTrain as an AI Safety Red Teamer testing conversational models with jailbreaks, prompt injections, and adversarial attacks. This contract, remote role pays $20–$22/hr, requires native English and Odia, and expects 20+ hours/week.

Generative AI & RLHF
Text
Remote · Worldwide
English, Odia
Part-time · Flexible
Expert level
Hourly · $20–$22/hr

Posted Jul 13, 2026