Skip to content
OpenTrain AIFor AI Companies

Bilingual LLM Safety Evaluator (French/English)

Join OpenTrain as a remote contractor to evaluate and red-team LLM outputs in French and English, focusing on safety, policy alignment, and adversarial case curation. This part-time role (20+ hrs/week) pays $24–$36/hr (typical $30/hr) and requires hands-on LLM red-teaming experience.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $24–$36/hr

$24–$36/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Apr 3, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the leading platform for building careers in AI training and data labeling. We hire contributors as contractors to do the human work that makes modern AI systems safer and more useful.

Working with OpenTrain means joining a fast-growing industry where people of many backgrounds shape how state-of-the-art models behave. Many roles are remote, flexible, and accessible without prior machine-learning credentials.

About AI Training and Safety Work

AI training (also called data labeling or human evaluation) is the human side of model development: people create, rate, and refine examples that teach AI how to respond. Safety evaluation and red teaming help prevent toxic, illegal, or otherwise harmful outputs before they reach users.

This role focuses on evaluating model responses, documenting adversarial patterns, and building robust labeling standards so models follow safety policies across French and English.

The Role

We are hiring a bilingual (French/English) LLM Safety Evaluator to score, annotate, and curate safety-focused evaluation content. You will produce red-team training cases, assess model outputs for policy alignment, and explain decisions in ambiguous or sensitive situations.

This is a contractor, part-time position (20+ hours/week), fully remote and open worldwide. You will work with text data and evaluation workflows (RLHF, red teaming, rating tasks).

  • Employment type: Contractor, Part-time
  • Time commitment: 20+ hours per week
  • Data type: Text — evaluation, RLHF, and red-team content

What You'll Do

Create, probe, and document adversarial prompts and model interactions in French and English to surface safety failures and edge-case behavior.

Score and annotate model outputs against written safety policies, explaining your decisions clearly for ambiguous or nuanced cases.

  • Curate red-team training cases across nuanced content domains (e.g., sexual content, self-harm, violence, hate, misinformation).
  • Produce detailed notes on adversarial patterns and safety boundary probes.
  • Apply written safety policies consistently and provide rationale that can be used to train labelers and improve model behavior.
  • Work with common AI tools (e.g., ChatGPT, Gemini, Perplexity) to run experiments and document findings.

Requirements

Candidates must meet the language, education, and experience requirements below exactly as stated.

This work involves regular exposure to explicit, toxic, violent, sexual, or psychologically disturbing content; you must be comfortable with that.

  • Near-native or native French proficiency (reading and writing).
  • Minimum C1 English proficiency (reading and writing).
  • Bachelor’s degree or higher in Communications, Linguistics, Psychology, Law/Policy, Security Studies, or equivalent professional experience.
  • Proven experience in Trust & Safety, content moderation, policy enforcement, risk operations, investigations, or safety evaluation.
  • Required hands-on LLM red teaming experience, including probing safety boundaries and documenting adversarial patterns.
  • Strong knowledge of safety categories: hate/harassment, sexual content, suicide/self-harm, violence, bias, illegal goods/services, malicious activities, malicious code, misinformation.
  • Ability to apply written safety policies consistently and explain decisions clearly in ambiguous cases.

Preferred Qualifications

These are helpful but not strictly required if you already meet the core requirements above.

Experience with annotation or AI data training workflows speeds onboarding and improves impact.

  • Prior experience in AI data training, annotation, or evaluation workflows.
  • Practical experience using tools like Perplexity, Gemini, ChatGPT, or similar systems to test and document model behavior.

Compensation, Schedule, and How It Works

Pay is hourly: typical rate $30 USD/hr with an allowable range of $24–$36 USD/hr. Payment is hourly as a contractor.

You will be asked to follow written guidelines and provide clear annotations and reasoning. Work is remote and flexible, but you must meet the 20+ hours/week expectation and deliver agreed tasks on time.

  • Payment type: Pay per hour (USD).
  • Hourly pay range: $24–$36 (typical $30/hr).
  • Location: Fully remote, worldwide.
  • Labeling focus: EVALUATION_RATING, RLHF, RED_TEAMING.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Bilingual LLM Safety Evaluator (Hebrew & English)

Join OpenTrain AI as a remote, part-time contractor reviewing and red-teaming LLM outputs in Hebrew and English to find safety failures and produce labeled evaluation data. $26–$38/hr, 20+ hours/week; your feedback will directly shape model safety.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $26–$38/hr

Posted Apr 3, 2026

French LLM Evaluator, Bilingual Linguist (BA Required)

Review and revise AI-generated French responses as a remote hourly contractor, rating outputs and producing corrected model answers; BA in linguistics/translation or a related field is required. Paid $16–$26.50/hr (typical $24/hr); worldwide and fully remote.

Generative AI & RLHF
Text
Remote · Worldwide
Flexible hours
Expert level
Hourly · $16–$26.5/hr

Posted Apr 3, 2026

Bilingual Arabic–English LLM Safety Evaluator

Join OpenTrain AI as a remote, part-time contractor to evaluate and label LLM outputs in Arabic and English, focusing on safety and adversarial behavior. This role requires 20+ hours/week and pays $15–$40/hr while you assess, annotate, and document unsafe or sensitive model responses.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $15–$40/hr

Posted Apr 3, 2026