Join OpenTrain AI as a contractor evaluating a model personalization feature in Arabic: design multi-turn prompts, compare side-by-side responses, and write clear rationales. 3-month contract at $15/hr with a flexible schedule (20+ hrs/week typical) and required PST overlap.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people start and grow careers in AI training and data labeling. We help freelancers find specialized AI training projects, build a unified portfolio, and turn annotation and evaluation work into a durable, remote career.
OpenTrain AI is the hiring and contracting organization for this role.
Work is fully remote and designed to be flexible for part-time contributors.
About AI training and personalization work
AI training (data labeling, annotation, and evaluation) is the human side of building intelligent systems. Contributors design prompts, judge model outputs, and provide the feedback that helps models personalize and behave more helpfully.
This project focuses on personalization: testing whether a model uses personal information naturally, stays grounded in facts, and produces useful, enjoyable responses.
Great for flexible, remote work that fits around other commitments.
No special corporate background required — attention to detail and clear writing matter most.
The role
You will work as an AI Quality Analyst (Personalization) — Arabic, evaluating a new personalization feature by designing prompts based on your own personal data and judging model responses for grounding, integration, and helpfulness.
This is a contractor position for a 3-month engagement at $15.00 USD per hour. The schedule is flexible but committed: at least 4 hours per day (up to 40 hours/week), with a typical expectation of 20+ hours/week; you must provide 4 hours of overlap with Pacific Standard Time (PST).
Contract type: Contractor, part-time possible up to full-time hours during the engagement.
Pay: $15/hour, paid by OpenTrain AI.
Language: High competence reading and writing Arabic is required.
Work is remote; you must use your own laptop/desktop and reliable internet.
What you'll do day to day
Your core work is prompt design and response evaluation. You will create multi-turn conversational prompts using your personal context, run side-by-side comparisons of model outputs, and write structured rationales explaining which response is better and why.
Design and execute 1–5 turn conversational prompts that require the model to use your personal information and experiences.
Evaluate responses for appropriate personalization relative to your intent.
Check grounding: identify unsupported claims or hallucinations in model outputs.
Assess integration: ensure the model uses personal data naturally without overnarration or robotic phrasing.
Rank two model responses side-by-side (SxS) and write clear, turn-referenced rationales.
Maintain data hygiene by deleting evaluation conversations after tasks are complete.
Requirements
All required skills come from the project brief. You must be able to work independently, follow evaluation guidelines closely, and communicate your judgments clearly in Arabic and (where asked) in English.
Fluent reading and writing ability in Arabic.
Experience designing creative multi-turn prompts for AI evaluation.
Strong analytical thinking and evaluation acumen for personalization concepts.
Ability to identify grounding errors and incorrect personalization.
Meticulous attention to detail for side-by-side comparisons and structured rationales.
Excellent written communication and ability to provide constructive feedback.
Reliable desktop or laptop and stable internet connection.
Availability to meet the schedule requirement including 4 hours overlap with PST.
Helpful background
The project prefers candidates with related study or experience but these are not strict requirements. Bring any annotation or moderation experience you have, plus analytical training from related fields.
BS/BA or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or another analytical field is a plus.
Prior experience in data annotation, AI quality evaluation, or content moderation is strongly preferred.
How the project works
You will receive evaluation tasks, create and run your prompts, and submit side-by-side rankings with turn-specific rationales. Follow data hygiene rules: delete or remove any evaluation conversations as instructed. Work is paid hourly and tracked according to OpenTrain AI procedures.
3-month contract paid at $15/hr; track hours as instructed by OpenTrain AI.
Expect to compose clear, concise written rationales referencing specific turns in each conversation.
All work is remote; you must be self-motivated and able to meet deadlines and overlap requirements.
Immediate 3-month contract evaluating Turkish-language personalization for a generative AI; $15/hr, minimum 30 hrs/week with 4 hours overlap with PST. Design multi-turn prompts, compare side-by-side responses, and write clear rationales to shape model behavior.
Evaluate a new personalization feature for a conversational AI in German: design multi-turn prompts, compare side-by-side responses, and write clear rationales. Contractor role, $15/hr, 3-month engagement, remote with required PST overlap.
Contract with OpenTrain as an AI Quality Analyst evaluating a conversational AI personalization feature: design multi-turn prompts, rate side-by-side text responses, and write defensible rationales. 3-month contract, remote, 20+ hrs/week with required PST overlap.