Skip to content
OpenTrain AIFor AI Companies

Arabic English AI Safety Content Evaluator

Evaluate AI-generated content in Arabic and English, identify unsafe or adversarial behavior, and provide clear feedback for safer language models. This fully remote contractor role offers $15-$40 per hour with a 20+ hour weekly commitment.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $15–$40/hr

$15–$40/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Apr 3, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this role, giving contributors a way to work directly on the human side of modern AI development.

Creating an OpenTrain account is free, and candidates can apply in minutes for opportunities that match their experience.

  • Fully remote contract work
  • Part-time schedule of 20+ hours per week
  • Hourly pay ranging from $15 to $40 USD
  • Intermediate-level opportunity

About AI Safety Evaluation

AI training depends on people who review model responses, identify risks, and explain how systems should improve. Safety evaluators help shape how large language models respond to sensitive, adversarial, and potentially harmful requests.

This work combines content evaluation, annotation, response writing, and human feedback. Your decisions can help reduce toxic, unsafe, misleading, or otherwise unwanted model behavior.

  • Contribute to the development of safer AI systems
  • Review and evaluate generated text
  • Apply written policies to complex and ambiguous cases
  • Work with sensitive material as part of routine evaluation

The Role

As an AI Safety Content Evaluator, you will review AI-generated responses and create safety-focused evaluation content in both English and Arabic. You will assess reasoning quality, annotate outputs for safety concerns, and provide expert feedback so responses are accurate, safe, and clearly explained.

The work may involve explicit, toxic, violent, sexual, or psychologically disturbing material. You will follow strict safety guidelines and documentation standards while evaluating content designed to test or expose unsafe model behavior.

  • Near-native or native Arabic reading and writing required
  • Minimum C1 English reading and writing proficiency required
  • 20+ hours per week
  • Hourly contractor engagement
  • Worldwide remote opportunity

What You'll Do

You will use critical judgment and advanced language skills to evaluate how language models respond to ordinary and adversarial prompts. Your feedback will help document unsafe behaviors and support consistent safety decisions across sensitive content categories.

  • Review AI-generated responses in Arabic and English
  • Assess reasoning quality and clarity
  • Annotate outputs for safety concerns
  • Identify adversarial prompts and unsafe model behaviors
  • Create safety-focused evaluation content
  • Provide clear, well-documented feedback
  • Evaluate content involving hate and harassment
  • Review sexual content, suicide and self-harm, violence, and psychologically disturbing material
  • Assess risks involving bias, illegal goods or services, malicious activities, malicious code, and misinformation
  • Apply written safety policies consistently in ambiguous cases

Requirements

This role requires demonstrated trust and safety judgment, hands-on LLM red teaming experience, and the ability to explain evaluation decisions clearly. A bachelor's degree or higher in a relevant field is required, or equivalent professional experience.

  • Bachelor's degree or higher in Communications, Linguistics, Psychology, Law or Policy, Security Studies, or a related field, or equivalent professional experience
  • Proven experience in Trust and Safety, content moderation, policy enforcement, risk operations, investigations, or safety evaluation
  • Hands-on LLM red teaming experience
  • Strong knowledge of the listed AI safety domains
  • Ability to apply written safety policies consistently
  • Ability to explain decisions clearly in ambiguous cases
  • Comfort reviewing explicit, toxic, violent, sexual, or psychologically disturbing content daily
  • Strong practical experience with Perplexity, Gemini, ChatGPT, or similar AI systems
  • Prior AI data training, annotation, or evaluation experience preferred

Who Should Apply

This opportunity is suited to bilingual Arabic and English professionals who can combine careful language review with policy-based risk assessment. It may be a strong fit for people with backgrounds in trust and safety, content moderation, investigations, risk operations, policy enforcement, or LLM red teaming.

Applicants should be comfortable making consistent judgments about difficult material, documenting their reasoning, and working with AI tools as part of an evaluation workflow.

  • Arabic and English bilingual professionals
  • Trust and safety specialists
  • Content moderators and policy enforcement professionals
  • Risk operations or investigations professionals
  • LLM red teamers and AI safety evaluators
  • Professionals with prior AI annotation or evaluation experience

How to Apply

Create a free OpenTrain account and apply through OpenTrain AI. Include your relevant trust and safety, LLM red teaming, language, and AI evaluation experience so your background can be considered for this remote contractor opportunity.

  • Apply through OpenTrain
  • Highlight Arabic and English proficiency
  • Describe relevant safety evaluation and red teaming experience
  • Showcase experience with AI systems and evaluation workflows

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Arabic AI Safety Evaluation Expert

Use Arabic fluency and cultural judgment to evaluate sensitive AI prompts, identify adversarial patterns, and improve model safety. This remote contractor role offers $28-$32 per hour with a default commitment of 7 hours per week.

Generative AI & RLHF
Text
Remote · Worldwide
Arabic, English
Part-time · Flexible
Entry level
Hourly · $28–$32/hr

Posted Sep 4, 2026

Arabic Scientific AI Safety Reviewer

Use advanced scientific expertise and Arabic fluency to write prompts, evaluate AI responses, and review sensitive dual-use content. This remote contract offers 7 hours per week at $38-$42 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
Arabic, English
Part-time · Flexible
Entry level
Hourly · $38–$42/hr

Posted Sep 5, 2026

Arabic AI Response Evaluator

Design Arabic multi-turn prompts and evaluate how naturally and accurately AI uses personal context. Join a remote, three-month contractor engagement paying $15 per hour through OpenTrain.

Generative AI & RLHF
Text
Remote · Worldwide
Arabic
Part-time · Flexible
Intermediate level
Hourly · $15/hr

Posted Jul 16, 2026