Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Teamer

Challenge frontier AI systems with adversarial prompts, uncover safety weaknesses, and document model behavior across high-risk topics. This expert contractor role pays $70 to $84 per hour for 20+ hours weekly.

OpenTrain AI

Generative AI & RLHF

Remote Hourly · $70–$84/hr

$70–$84/hr

Compensation

33 countries

Eligibility

Expert

Experience

Jul 17, 2026

Posted

Open to applicants in

United States Denmark Estonia Finland Ireland Latvia Lithuania Norway Sweden Austria Belgium France Germany Netherlands Switzerland United Kingdom Albania Bosnia & Herzegovina Croatia Greece Italy Malta Portugal Serbia Slovenia Spain Bulgaria Czechia Hungary Moldova Poland Romania Slovakia

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build a credible AI training profile, and apply for work that matches their experience.

Creating an OpenTrain account is free. Your profile can help you present your experience in AI safety, red teaming, trust and safety, and related fields as you pursue a longer-term career in AI training.

About AI Safety Red Teaming

AI training is the human work behind modern artificial intelligence. Contributors write prompts, evaluate model responses, identify failures, and provide structured feedback that helps AI systems become more reliable and safer.

Red teaming applies this process adversarially. Instead of testing only expected behavior, you will probe model boundaries, surface vulnerabilities, and examine how systems respond to ambiguous, sensitive, or high-risk scenarios.

The Role

OpenTrain is recruiting an AI Safety Red Teamer to perform adversarial testing of frontier AI systems. The work combines structured evaluation, safety judgment, prompt design, and clear documentation to support model robustness and alignment.

This is an expert-level, part-time contractor opportunity requiring 20 or more hours per week. The role is available to candidates in the listed countries and requires English-language work.

  • Payment: $70 to $84 USD per hour
  • Experience level: Expert
  • Workload: 20+ hours per week
  • Engagement: Contractor and part-time
  • Language: English

What You'll Do

You will design challenging prompts and evaluate model behavior across complex, high-risk, and ambiguous topics. Findings must be recorded clearly so researchers and safety teams can understand the issue, assess its significance, and strengthen safeguards.

Your evaluations may cover cyber, biosecurity, fraud, political content, scientific safety, and related sensitive domains.

  • Create adversarial prompts that probe model boundaries and expose vulnerabilities.
  • Test for jailbreaks, unsafe behavior, hallucinations, misinformation, policy failures, and other reliability concerns.
  • Evaluate responses across cyber, biosecurity, fraud, political content, scientific safety, and related sensitive areas.
  • Record findings clearly and contribute to safety benchmarking and red-teaming reports.
  • Collaborate with AI researchers and safety teams to improve model safeguards.

Required Qualifications

A bachelor's degree or higher in computer science, cybersecurity, journalism, communications, psychology, biology, chemistry, public policy, or a related discipline is required. You must also have at least five years of professional experience in AI safety, AI red teaming, trust and safety, cybersecurity, investigative journalism, life sciences, or a related field.

Demonstrated experience designing adversarial prompts or evaluating frontier AI systems is required. Strong analytical reasoning, prompt design, and written communication skills are essential for this work.

  • Bachelor's degree or higher in a relevant discipline.
  • At least five years of professional experience in a related field.
  • Experience designing adversarial prompts and testing frontier AI systems.
  • Ability to identify jailbreaks, hallucinations, unsafe behavior, and policy failures.
  • Sound judgment across cyber, biosecurity, misinformation, fraud, or political-content scenarios.
  • Strong analytical reasoning and written documentation skills.

Helpful Background

Experience with AI red teaming, reinforcement learning from human feedback, supervised fine-tuning, AI alignment, trust and safety, jailbreak testing, prompt engineering, or adversarial evaluation methodologies is valuable.

Specialized knowledge of cyber, biosecurity, political content, misinformation, or scientific safety can help you assess difficult cases with appropriate context and care.

  • AI red teaming or adversarial evaluation methodologies
  • RLHF, SFT, or AI alignment
  • Trust and safety or jailbreak testing
  • Prompt engineering
  • Cyber, biosecurity, political content, misinformation, or scientific safety expertise

Why This Work Matters

Every major AI system depends on human contributors who prepare examples, assess outputs, and identify weaknesses. By testing how frontier models behave under pressure, red teamers help shape safer and more dependable AI systems.

AI training is a fast-growing field that can offer flexible, remote work for people with specialized expertise. This role gives experienced professionals a direct way to apply their judgment to cutting-edge AI development.

Eligible Locations

This opportunity is open to candidates located in the United States, Denmark, Estonia, Finland, Ireland, Latvia, Lithuania, Norway, Sweden, Austria, Belgium, France, Germany, the Netherlands, Switzerland, the United Kingdom, Albania, Bosnia and Herzegovina, Croatia, Greece, Italy, Malta, Portugal, Serbia, Slovenia, Spain, Bulgaria, Czechia, Hungary, Moldova, Poland, Romania, or Slovakia.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Red-Teaming Quality Assurance Lead

Lead quality assurance for AI red-teaming and safety evaluation projects, reviewing adversarial prompts, risk analyses, and contributor work. This remote U.S. contract role offers up to $100 per hour and requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $100/hr

Posted Jul 9, 2026

AI Safety Red Team Expert

Probe conversational AI for jailbreaks, prompt injections, bias exploitation, and manipulation as an expert red team contractor. Work worldwide for $48 to $62 per hour, 20+ hours weekly, using English and Danish.

Generative AI & RLHF
Text
Remote · Worldwide
English, Danish
Part-time · Flexible
Expert level
Hourly · $48–$62/hr

Posted Jul 30, 2026

AI Safety Red Team Expert English Indonesian

Probe conversational AI models and agents for vulnerabilities using adversarial testing, jailbreaks, prompt injections, and multi-turn manipulation. This remote, part-time expert contract pays $17 to $25 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
Indonesian, English
Part-time · Flexible
Expert level
Hourly · $17–$25/hr

Posted Jul 30, 2026