Skip to content
OpenTrain AIFor AI Companies

AI Safety Red Team And Policy Evaluator

Design adversarial prompts, review image-edit requests, and evaluate AI model safety against detailed policies. This worldwide, English-language contract requires 20 or more hours each week.

Apply now
OpenTrain AI

Generative AI & RLHF

100% Remote Fixed price

Fixed price

Compensation

Worldwide

Eligibility

Entry

Experience

Sep 26, 2026

Posted

Open worldwide

The Work

You will help evaluate and improve safety behavior in large language models. The work combines red-team testing, policy classification, output review, and clear written explanations that can be used to improve models.

  • Create adversarial prompts that test the limits of AI safety filters.
  • Review single-turn image-edit requests and outputs using a safety taxonomy.
  • Classify content involving areas such as self-harm, violence, hate speech, sexual content, minors, and identity depiction.
  • Apply safety rules consistently to unclear and borderline cases.
  • Write evaluation rubrics and rationales that clearly distinguish between model outputs.
  • Document bypassed safeguards and subtle policy violations.
  • Optionally compare and rank multiple model outputs for the same request.
  • Identify unclear or conflicting guidelines and suggest improvements.

What It Pays And Takes

This is a part-time contractor role with fixed-price payment. The source listing does not provide a fixed payment amount.

  • Time requirement: 20 or more hours per week.
  • Location: Worldwide.
  • Language: English.
  • Work setup: You must provide your own desktop or laptop and reliable internet connection.
  • You need strong judgment when applying policy criteria to complex or ambiguous content.
  • You need experience with red teaming, prompt engineering, or designing challenge prompts for AI safety filters.
  • You need a solid understanding of trust and safety principles for large language models.
  • You need precise writing skills for self-contained rubrics and concise rationales.
  • Experience with RLHF workflows or data annotation is a significant plus.
  • A degree or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or a related field is helpful.
  • Previous work in content moderation, policy analysis, AI safety evaluation, or similar work is strongly preferred.

How It Works

Apply on OpenTrain with your resume, then complete the application on the hiring site.

About AI Training Work

AI training is the human work behind systems that learn from examples, including reviewing model responses, classifying content, and testing safety behavior. People with strong judgment and relevant policy or technical knowledge help create reliable feedback that improves how AI systems respond.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Red-Teaming Quality Assurance Lead

Review AI safety evaluations, adversarial prompts, and red-team submissions as a remote quality assurance lead. This US contract role pays up to $100 per hour and requires 20+ hours weekly.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $100/hr

Posted Jul 9, 2026

AI Safety LLM Evaluator, French and English

Review AI-generated responses and create French-English safety evaluations as a fully remote contractor. Score outputs, write red-team cases, and help reduce unsafe model behavior for $24 to $36 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
French
Part-time · Flexible
Intermediate level
Hourly · $24–$36/hr

Posted Apr 3, 2026

AI Red Team Engineer for LLM Security

Work remotely as an AI Red Team Engineer testing LLMs, agents, and RAG systems for security weaknesses. This contract role pays $40 per hour and requires C1 English plus hands-on cybersecurity experience.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level
Hourly · $40/hr

Posted Oct 6, 2025