Skip to content
OpenTrain AIFor AI Companies

Bilingual AI Safety Data Evaluator

Evaluate AI-generated content and safety decisions in Spanish and English as a remote contractor earning $14-$24 per hour. Use red-teaming, policy judgment, and clear analytical feedback to help improve safer AI systems.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $14–$24/hr

$14–$24/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Apr 3, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is hiring a Bilingual AI Safety Data Evaluator for remote, hourly contract work. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people contribute to the development of cutting-edge AI systems.

In this role, you will work at the human side of AI development by reviewing model outputs, identifying safety risks, and creating feedback that helps AI behave more accurately, safely, and responsibly.

  • Remote worldwide opportunity
  • Independent contractor engagement
  • Hourly pay of $14-$24 USD, with a listed rate of $20 USD per hour
  • Spanish and English language work

The Role

You will review AI-generated content and evaluate the quality of safety and reasoning decisions in both Spanish and English. Your annotations and feedback will help assess whether model outputs are accurate, safe, logical, and well explained.

The work includes nuanced policy judgments, multilingual content review, and adversarial testing. You may encounter explicit, toxic, violent, sexual, or psychologically disturbing material as part of your daily responsibilities.

  • Evaluate AI safety and reasoning outputs
  • Apply policies consistently across Spanish and English content
  • Identify risky, unsafe, or poorly reasoned responses
  • Provide clear and reproducible rationales for decisions

What You'll Do

You will combine content evaluation, safety data quality checks, and LLM red-teaming to find edge cases that may be missed by standard reviews. Your work will include recommending mitigations and preserving the meaning, severity, and intent of content across languages.

Strong analytical writing is central to the role. Evaluations should clearly explain why a response is safe, unsafe, inaccurate, biased, or otherwise inconsistent with applicable policy guidance.

  • Label and evaluate AI-generated text
  • Quality-check safety data and annotations
  • Red-team AI systems for edge cases and adversarial behavior
  • Review risks involving hate and harassment, sexual content, self-harm, violence, and bias
  • Assess content related to illegal goods or services, malicious activities, malicious code, and misinformation
  • Recommend mitigations for identified safety issues
  • Review multilingual and cross-cultural content in Spanish and English

Requirements

This is an intermediate-level role for an experienced safety, moderation, risk, or policy professional. Candidates should be able to make careful judgments on sensitive material, explain those judgments in writing, and apply guidelines consistently across languages and cultures.

  • Near-native or native Spanish proficiency in reading and writing
  • Minimum C1 English proficiency in reading and writing
  • Bachelor’s degree or higher in Communications, Linguistics, Psychology, Law or Policy, Security Studies, or a related field, or equivalent professional experience
  • At least 5 years of professional experience in Trust & Safety, content moderation, policy operations, risk, compliance, investigations, or related safety work
  • Proven LLM red-teaming or adversarial testing experience
  • Experience identifying edge cases and recommending mitigations
  • Strong knowledge of relevant AI safety domains, including hate, self-harm, violence, bias, illegal activity, malicious code, and misinformation
  • Experience applying policy guidelines consistently across Spanish and English content
  • Strong analytical writing skills and clear, reproducible decision rationales
  • Comfort reviewing explicit, toxic, violent, sexual, or psychologically disturbing content

Preferred Experience

Localization or translation experience is preferred. This background can help you assess whether an AI response preserves meaning, severity, and intent when reviewed across Spanish and English.

  • Localization experience
  • Translation experience
  • Experience with multilingual or cross-cultural safety content
  • Familiarity with nuanced policy interpretation and safety operations

Why AI Safety Work Matters

AI training and data evaluation are the human processes behind modern AI systems. Reviewers assess examples, rate model responses, and provide feedback that influences how systems understand language, follow safety standards, and respond to people.

This remote work offers a chance to contribute directly to the development of state-of-the-art AI while using specialized experience in Trust & Safety, content moderation, and multilingual policy evaluation.

  • Help shape how AI systems handle sensitive and high-risk content
  • Apply professional safety expertise to emerging technology
  • Work remotely from anywhere with an internet connection
  • Contribute to a fast-growing field at the intersection of language and AI

How to Apply

Create a free OpenTrain account and apply in minutes. Your experience with AI safety evaluation, LLM red-teaming, multilingual policy work, and sensitive content review will be especially relevant to this opportunity.

  • Review the role details and requirements
  • Create or update your OpenTrain profile
  • Highlight Spanish and English proficiency and relevant safety experience
  • Submit your application through OpenTrain

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Chinese Bilingual STEM AI Safety Evaluator

Use Chinese and scientific expertise to evaluate AI responses for accuracy, helpfulness, and safe handling of sensitive STEM topics. This flexible contractor role pays $68 to $72 per hour for fewer than 20 hours per week.

Generative AI & RLHF
Text
Remote · Worldwide
Chinese, English
Part-time · Flexible
Entry level
Hourly · $68–$72/hr

Posted Sep 6, 2026

AI Safety LLM Evaluator, French and English

Work remotely as a French and English AI Safety LLM Evaluator, reviewing model responses, red-teaming safety boundaries, and creating evaluation data. Earn $24 to $36 per hour while helping improve safer AI systems.

Generative AI & RLHF
Text
Remote · Worldwide
French
Part-time · Flexible
Intermediate level
Hourly · $24–$36/hr

Posted Apr 3, 2026

Spanish AI Safety Content Evaluator

Use native-level Spanish, cultural judgment, and careful reasoning to evaluate sensitive AI content and help improve model safety. This remote contractor role offers $43-$47 per hour with a default commitment of 7 hours weekly.

Generative AI & RLHF
Text
Remote · Worldwide
Spanish, English
Part-time · Flexible
Entry level
Hourly · $43–$47/hr

Posted Sep 4, 2026