Evaluate and shape AI personalization in Indonesian — $15/hr, contractor work with a minimum of 20 hours/week; ideal candidates can commit 30+ hours with PST overlap. Design multi-turn prompts, rate model outputs, write clear rationales, and verify grounding.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We help people discover opportunities, build a unified AI-training portfolio, and grow durable freelance careers in a fast-growing field.
OpenTrain is the hiring and contracting organization for this project. Creating an OpenTrain account is free.
About AI training and personalization
AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. Personalization work tests how models use personal context to provide relevant, helpful responses; your judgments directly shape how AI understands and applies user information.
This role focuses on evaluating personalized conversational behavior: designing prompts that use your personal context, checking whether the model grounds its answers in that context, and measuring integration and usefulness across turns.
The role
We are hiring AI Quality Analysts to evaluate a personalization feature in Indonesian. This is remote contract work where contributors design multi-turn prompts, compare model outputs, and write defensible rationales explaining their judgments.
Position type: Contractor, part-time (contractor relationship through OpenTrain).
Pay: $15.00 USD per hour.
Minimum commitment: 20+ hours/week; ideal candidates can commit 30+ hours/week with overlap during Pacific Time (PST).
Work language: Indonesian (read and write at a high level).
Data type: Text; label type: evaluation/rating (side-by-side comparisons, structured rationales).
Worldwide applicants welcome — you must be able to work the required hours and overlap with PST when requested.
What you'll do
Your core responsibility is to test and assess how well the model uses personal information across multi-turn conversations. You will create scenarios that require the model to reference personal context, judge response quality, and produce detailed rationales.
Design creative, multi-turn conversational prompts that rely on your personal information and experiences.
Evaluate model responses for Grounding (uses correct personal info), Integration (applies context across turns), and Helpfulness.
Perform side-by-side comparisons of two model responses and rank them.
Write clear, defensible rationales tied to specific conversation turns and evidence.
Extract and verify debug information to confirm proper data source usage when requested.
Maintain data hygiene by deleting evaluation conversations after completion.
Requirements
You must be comfortable working independently, thinking critically about nuanced conversational behavior, and producing structured written evaluations in Indonesian.
Fluency in reading and writing Indonesian at a high level.
Availability for at least 20 hours/week; preference for candidates able to work 30+ hours/week with PST overlap.
Strong analytical thinking and the ability to evaluate nuanced AI responses.
Experience in creative prompt engineering or similar test-crafting work.
Proven ability to perform careful side-by-side comparisons and produce meticulous, evidence-based rationales.
Excellent written communication skills in Indonesian and strong attention to detail.
Self-motivated, reliable, and able to work remotely as a contractor.
Helpful background
These qualifications are helpful but not strictly required. Candidates with related education or prior annotation experience will be competitive.
BS/BA or equivalent in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or related fields.
Prior experience in data annotation, AI quality evaluation, content moderation, or similar human-evaluation roles.
Remote contractor role evaluating a conversational AI personalization feature in Thai by designing multi-turn prompts, comparing side-by-side responses, and writing structured rationales. 3-month contract at $15/hr with a minimum 4 hours/day and 4-hour overlap to PST.
Join OpenTrain to evaluate a new personalization feature in Japanese: design multi-turn prompts, compare paired model responses, and write clear rationales in a remote, 3-month contractor role paying $15/hr with PST overlap.
Contract with OpenTrain as an AI Quality Analyst evaluating a conversational AI personalization feature: design multi-turn prompts, rate side-by-side text responses, and write defensible rationales. 3-month contract, remote, 20+ hrs/week with required PST overlap.