Remote contractor role evaluating a conversational AI personalization feature in Thai by designing multi-turn prompts, comparing side-by-side responses, and writing structured rationales. 3-month contract at $15/hr with a minimum 4 hours/day and 4-hour overlap to PST.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. Creating an OpenTrain account is free — build a profile, discover projects across the industry, and start applying quickly to contract opportunities that let you grow a durable freelance career in AI training.
Why AI training work matters
AI training (data labeling, annotation, and human evaluation) is the human side of building intelligent systems. Contributors design prompts, evaluate outputs, and teach models to behave better — work that is remote, flexible, and accessible to many skill levels.
This project focuses on personalization quality, giving you direct influence over how conversational AI uses personal context to produce helpful, grounded responses.
100% remote: work from anywhere with a computer and internet
Flexible hours: many contributors set a schedule that fits other commitments
Accessible: subject-matter expertise helps, but many roles need strong language and analytical skills
The role
You will be contracted by OpenTrain AI as an AI Quality Analyst focused on Thai-language personalization evaluation. This hands-on role asks you to design multi-turn conversational prompts that require the model to draw on your personal context, then assess how well the model integrates that context into its responses.
This is a practical evaluation role for people who enjoy creative prompt design, detailed analytical review, and clear written reasoning. You will compare pairs of model responses and produce structured, turn-referenced rationales for your rankings.
What you'll do
Design and run multi-turn conversational prompts (1–5 turns) that require the model to use your personal information and experiences (e.g., email, search, watch history).
Evaluate model outputs for Grounding, Integration, and Helpfulness, flagging incorrect personalization, poor inferences, and forced connections.
Stack-rank two model responses side-by-side (SxS) to identify which is more helpful, easier to use, and more enjoyable.
Write clear, structured rationales for your rankings that reference specific turn numbers and concrete examples.
Maintain data hygiene by deleting evaluation conversations so your future chat history remains unpolluted.
Requirements
Thai proficiency: able to read and write Thai at a high level.
Exceptional analytical thinking to evaluate nuanced or ambiguous personalization behavior.
Experience designing creative multi-turn prompts that use personal context.
Strong evaluation acumen: able to identify incorrect personalization, poor inferences, and forced connections.
Meticulous attention to detail when reviewing side-by-side model responses.
Excellent written communication for concise, structured rationales.
Self-motivated and able to work independently in a remote contractor setting.
Desktop or laptop with a reliable internet connection.
Availability: at least 4 hours per day with 4 hours of overlap to Pacific Standard Time (PST); typical weekly workload 20+ hours, up to 40 hours/week.
Engagement length: 3 months. Pay rate: $15 per hour.
Helpful background (not required)
BS/BA or equivalent experience in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field.
Prior experience in data annotation, AI quality evaluation, or content moderation.
How to apply and next steps
OpenTrain AI is the contracting organization for this role. To apply, create a free OpenTrain account, complete your profile, and submit your application. Selected contributors will receive instructions and any role-specific orientation materials.
If you enjoy shaping how conversational AI uses personal context and can work in Thai with strong analytical rigor, this project is a great way to build experience in generative AI evaluation work.
Evaluate and shape AI personalization in Indonesian — $15/hr, contractor work with a minimum of 20 hours/week; ideal candidates can commit 30+ hours with PST overlap. Design multi-turn prompts, rate model outputs, write clear rationales, and verify grounding.
Join OpenTrain as a contractor evaluating a conversational AI personalization feature in Vietnamese. Earn $15/hr on a 3-month, fully remote engagement designing multi-turn prompts, rating responses side-by-side, and writing clear rationales with a required PST overlap.
Lead Thai-language QA for AI-generated content in a remote contractor role paying $25–$30 USD/hr. Review outputs, coach trainers/QAs, maintain Thai style guides, and help scale QA workflows with 20+ hours/week.