Join OpenTrain AI as a contract AI Quality Analyst (Dutch) to design multi-turn, personalization-based prompts and evaluate model responses for grounding, integration, and helpfulness; 1-month engagement at $20/hr with required PST overlap. Remote, flexible, entry-level.
Generative AI & RLHF
100% Remote Hourly · $20/hr
$20/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people start and grow careers in AI training and data labeling. We connect skilled contributors with short-term projects, help you build a unified AI training portfolio, and support a durable freelance career in this fast-growing industry.
Why this AI training work matters
AI training (data labeling and human evaluation) is the human side of building intelligent systems — people create and review examples that teach models how to respond, personalize, and behave. This work is highly flexible, remote-friendly, and places contributors on the cutting edge of how AI learns.
The role
OpenTrain AI is hiring an AI Quality Analyst to evaluate a new personalization feature. You will design creative, short multi-turn conversations that rely on your own personal context and rigorously assess how the model uses that context to respond.
This is a contractor, part-time assignment for one month. The project requires a minimum daily commitment and schedule overlap with Pacific Time; the offered rate is $20 per hour.
Engagement length: 1 month
Employment type: Contractor, Part-time
Rate: $20.00 USD per hour
Remote: worldwide candidates accepted
Language requirement: high-level Dutch (read/write)
What you'll do day-to-day
You will create short (1–5 turn) conversational prompts that require the model to use personal context, evaluate model outputs, and perform direct comparisons between responses.
Design and execute multi-turn conversational prompts based on your personal context or connected services.
Evaluate responses for Grounding issues (is the response supported by the personal context?), Integration quality (does the model weave personal details naturally?), and overall Helpfulness.
Perform side-by-side (SxS) comparisons: rank two model responses and write clear, defensible rationales for your choices.
Maintain strict data hygiene by deleting evaluation conversations after each session.
Requirements
You must be able to read and write Dutch at a high level and demonstrate strong analytical skills and careful attention to subtle differences in natural language responses.
Dutch proficiency: high reading and writing competency (required).
Experience designing creative, multi-turn conversational prompts that use personal context.
Ability to evaluate AI responses for Grounding, Integration, and Helpfulness.
Excellent written communication for producing clear, structured rationales.
Meticulous attention to detail and strong analytical thinking.
Helpful background (not required)
Candidates with related education or prior experience in evaluation roles will be at an advantage, though this is an entry-level opportunity for people with the right skills and mindset.
BS/BA or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or another analytical field.
Previous work in data annotation, AI quality evaluation, content moderation, or similar roles is helpful but not mandatory.
Schedule, payment, and next steps
This project requires a 1-month commitment. The client expects at least 4 hours per day and up to 40 hours per week, with a required 4-hour overlap with Pacific Time. The job listing indicates an overall time expectation of 20+ hours/week; please ensure you can meet the overlap requirement.
To apply, create or sign into your OpenTrain profile, confirm your availability and Dutch proficiency, and submit. OpenTrain is the hiring organization for this role and will manage contracting and payment.
Pay: $20.00 USD per hour, paid to contractors.
Time commitment: minimum 4 hours/day; up to 40 hours/week; overlap: 4 hours with PST.
Evaluate a new personalization feature for a conversational AI in German: design multi-turn prompts, compare side-by-side responses, and write clear rationales. Contractor role, $15/hr, 3-month engagement, remote with required PST overlap.
Evaluate and shape AI personalization in Indonesian — $15/hr, contractor work with a minimum of 20 hours/week; ideal candidates can commit 30+ hours with PST overlap. Design multi-turn prompts, rate model outputs, write clear rationales, and verify grounding.
Contract with OpenTrain as an AI Quality Analyst evaluating a conversational AI personalization feature: design multi-turn prompts, rate side-by-side text responses, and write defensible rationales. 3-month contract, remote, 20+ hrs/week with required PST overlap.