Evaluate a new personalization feature for a conversational AI in German: design multi-turn prompts, compare side-by-side responses, and write clear rationales. Contractor role, $15/hr, 3-month engagement, remote with required PST overlap.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 15, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow careers teaching AI by bringing specialized projects, tools, and a profile system together so contributors can discover work, track progress, and build a durable freelance career.
Why work in AI training
AI training (data labeling and human feedback work) is the human layer behind modern AI systems — people create, evaluate, and refine the examples models learn from. It’s often 100% remote, flexible, accessible without advanced credentials, and a direct way to shape how state-of-the-art AI behaves.
Work remotely from anywhere with a computer or phone and an internet connection.
Flexible part-time options let you fit work around studies, other jobs, or family.
Many projects require language fluency or attention to detail rather than formal degrees.
The role: AI Quality Analyst (Personalization) — German
You will evaluate how well a conversational AI personalizes responses using information from past conversations, Gmail, Search, and YouTube activity. The role blends creative prompt design with careful analytical evaluation of model output and requires high-quality written rationales in German.
This is a contractor, part-time engagement with a defined duration and specific schedule overlap requirements (see Compensation & Schedule).
Focus language: German — high competence reading and writing required.
Work type: remote contractor, contributor to an AI personalization evaluation project.
What you'll do day-to-day
You will design multi-turn conversational prompts from your own experiences and then evaluate how the model uses personal data to respond. Expect to compare outputs side-by-side and provide clear, defensible reasoning for which response is more helpful and why.
Design and run multi-turn prompts (typically 1–5 turns) that require the AI to use personal information.
Evaluate responses for correct application of personalization versus flawed inferences or hallucinations.
Check grounding and evidence: ensure claims about you are supported and not invented.
Assess integration quality so personal data reads naturally and not overnarrated or awkward.
Stack-rank two model responses side-by-side for helpfulness, usability, and enjoyability.
Write clear, structured rationales that reference exact points in the conversation.
Maintain strict data hygiene by deleting evaluation conversations after work.
Requirements
All requirements below are part of the selection criteria. You must have the tools and abilities to work independently and produce clear German-language evaluations.
German proficiency: strong reading and writing skills in German (required).
Creative prompt engineering experience for conversational AI evaluation.
Exceptional analytical thinking and nuance in evaluating AI responses.
Ability to evaluate grounding, integration quality, and subtle differences in side-by-side outputs.
Excellent written communication for concise, defensible rationales in German.
Self-motivated, able to work remotely with a desktop or laptop and reliable internet connection.
Helpful background
These are not required but will strengthen your application and help you ramp faster.
BS/BA or equivalent in Linguistics, Computer Science, Policy, Journalism, or similar fields.
Previous experience in data annotation, AI quality evaluation, or content moderation.
Compensation, schedule, and engagement
Pay, time commitment, and engagement length are fixed for this project and must be accepted as stated.
Hourly rate: $15 USD per hour.
Typical commitment: 20+ hours/week; expect at least 4 hours per day (up to 40 hours/week) with 4 hours overlap with Pacific Standard Time (PST).
Engagement length: 3 months as a contractor on a remote, global team.
How the process works
Apply through OpenTrain; selected contributors will receive onboarding instructions and evaluation tasks. You will be asked to follow evaluation guidelines closely and to delete conversational data after completing each task to maintain data hygiene.
Onboarding will cover evaluation rubrics, prompt design tips, and data hygiene rules.
Tasks involve designing prompts from your experience, running evaluations, and submitting side-by-side comparisons with written rationales in German.
Contract with OpenTrain as an AI Quality Analyst evaluating a conversational AI personalization feature: design multi-turn prompts, rate side-by-side text responses, and write defensible rationales. 3-month contract, remote, 20+ hrs/week with required PST overlap.
Join OpenTrain to evaluate a new personalization feature in Japanese: design multi-turn prompts, compare paired model responses, and write clear rationales in a remote, 3-month contractor role paying $15/hr with PST overlap.
Evaluate a personalization feature in Polish by designing short multi-turn prompts, comparing paired model responses, and writing concise, defensible quality rationales. Contractor role, $20/hr, remote, ~4 hours/day with 4-hour overlap with PST for a 1-month engagement.