Join OpenTrain to evaluate a new personalization feature in Japanese: design multi-turn prompts, compare paired model responses, and write clear rationales in a remote, 3-month contractor role paying $15/hr with PST overlap.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We help freelancers discover projects, collect credible proof-of-work, and grow long-term, flexible careers teaching AI.
This role is hired and managed by OpenTrain AI as part of our ongoing work to improve generative models through human evaluation.
About AI training and personalization work
AI training (data labeling and human evaluation) is the human side of building modern AI: people create, test, and judge examples that teach models to behave usefully and safely.
Personalization evaluation asks whether an assistant correctly and responsibly uses a real user's personal data to give helpful, context-aware responses — your judgments directly shape how assistants integrate personal information.
The role
You will be an AI Quality Analyst focused on Japanese-language personalization evaluation for a new assistant feature. This is a remote, 3-month contractor engagement where you design prompts using your own account data, rate paired responses, and write structured rationales.
Engagement type: contractor, part-time.
Language: Japanese (reading and writing required).
Focus: multi-turn conversational prompts and side-by-side model comparisons.
What you'll do
Your daily work combines creative prompt design with careful, evidence-based evaluation of model outputs. You will compare two model responses side-by-side and explain which is better and why.
Design and execute multi-turn conversational prompts (typically 1–5 turns) that require personalization from your own account data.
Evaluate responses on Grounding, Integration, and Helpfulness.
Stack-rank two model outputs side-by-side and write clear, defensible rationales referencing specific conversation points.
Identify incorrect personalization, poor inferences, and issues with naturalness or tone.
Maintain strict data hygiene by deleting evaluation conversations after each session.
Requirements
All listed requirements come from the project description and are necessary to perform the work as specified.
Fluent Japanese reading and writing is required.
Exceptional analytical thinking and attention to subtle differences in natural language.
Creative prompt engineering experience using personal context.
Strong written communication for concise, structured rationales.
BS/BA degree or equivalent experience in a relevant field (Policy, Law, Ethics, Linguistics, Journalism, Computer Science, etc.).
Experience in data annotation, AI quality evaluation, or content moderation is strongly preferred.
Self-motivated and able to work independently in a remote setting.
Schedule, pay, and logistics
This is a 3-month contractor role with part-time hours and a fixed hourly rate. You must overlap with Pacific Standard Time for coordination.
Commitment: at least 4 hours per day, up to 20 hours per week.
Required overlap: 4 hours per day with PST.
Rate: $15 USD per hour, paid to contractors.
Engagement length: 3 months.
Work is remote and may be done from many countries; you must be able to work the required PST overlap.
Contract with OpenTrain as an AI Quality Analyst evaluating a conversational AI personalization feature: design multi-turn prompts, rate side-by-side text responses, and write defensible rationales. 3-month contract, remote, 20+ hrs/week with required PST overlap.
Evaluate a new personalization feature for a conversational AI in German: design multi-turn prompts, compare side-by-side responses, and write clear rationales. Contractor role, $15/hr, 3-month engagement, remote with required PST overlap.
Lead QA for Japanese AI-generated content on OpenTrain: review outputs, coach trainers and QAs, maintain style guides, and help scale quality processes. Contract, 20+ hrs/week, $55/hr, remote within Japan.