Evaluate personalization for a leading generative AI in Hindi: design multi-turn prompts, compare and rate responses, and write clear rationales. Remote contractor role, 20+ hrs/week, $15/hr, requires strong Hindi reading/writing and overlap with PST.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect contributors with contract work that shapes how modern AI systems behave and help people build a durable freelance career in this fast-growing field.
Work 100% remotely from anywhere with a computer or phone and an internet connection.
Flexible part-time and contractor opportunities that fit around other commitments.
Build a unified AI training portfolio and grow into higher-paying, specialist projects over time.
About AI training and personalization work
AI training (also called data labeling, annotation, or human feedback) is the human side of building intelligent systems. Contributors create and review examples—like prompts and responses—that teach models how to behave.
Personalization evaluation focuses on how a model uses personal information to give grounded, helpful, and context-aware replies. Your judgments directly influence model safety, reliability, and user experience.
This role focuses on Hindi-language evaluation of personalization behavior.
Tasks are text-based and involve designing prompts, rating outputs, and recording clear rationales.
The role
You will be an AI Quality Analyst working on a personalization evaluation project for a leading generative model. As a contractor, you will design multi-turn conversational prompts, evaluate how the model uses personal information, and produce structured feedback to improve model behavior.
Position: Contractor, part-time (entry level).
Time commitment: 20+ hours per week; requires full-time availability with at least 4 hours overlap with Pacific Time (PST).
Pay: $15 USD per hour.
What you'll do day-to-day
Work through text evaluation tasks focused on personalization. Each task asks you to create or use multi-turn prompts and judge paired model responses for quality and correct use of personal data.
Design and execute 1–5 turn conversational prompts to test personalization.
Evaluate responses for grounding, integration of personal details, and overall helpfulness.
Stack-rank two model responses side-by-side and select the better reply.
Write clear, structured rationales that reference specific conversation turns.
Extract and verify debug information to confirm the model used proper data sources.
Maintain data hygiene by deleting evaluation conversations when required.
Requirements
We need dependable evaluators with strong Hindi language skills and a structured, analytical approach to rating model outputs. Preserve accuracy and clarity in every rationale you write.
Fluent Hindi reading and writing at a high level (required).
Analytical and critical thinking skills; attention to detail.
Experience with prompt engineering and designing multi-turn conversations.
Excellent written communication in Hindi and ability to produce clear rationales.
Availability for 20+ hours/week and at least 4 hours overlap with PST.
Employment type: contractor (part-time).
Helpful background (not required)
Candidates with related experience often ramp up faster, but entry-level applicants who meet the core requirements are encouraged to apply.
BS/BA or equivalent education in a relevant field.
Experience in data annotation, AI quality evaluation, or content moderation.
How it works and how to apply
OpenTrain manages contracting, onboarding, and assignment delivery. You will complete training tasks and calibration checks before receiving paid evaluation work.
To apply, create a free OpenTrain account, complete your profile, and submit your application for this role. Successful candidates will complete brief qualification tests and onboarding modules before beginning live tasks.
Work is entirely remote and text-based; label type: evaluation/rating.
You will be asked to demonstrate Hindi reading/writing and prompt design in a qualification test.
Payments are hourly at $15 USD; contractors are paid per OpenTrain procedures.
Evaluate and shape AI personalization in Indonesian — $15/hr, contractor work with a minimum of 20 hours/week; ideal candidates can commit 30+ hours with PST overlap. Design multi-turn prompts, rate model outputs, write clear rationales, and verify grounding.
Join OpenTrain to evaluate a new personalization feature in Japanese: design multi-turn prompts, compare paired model responses, and write clear rationales in a remote, 3-month contractor role paying $15/hr with PST overlap.
Contract with OpenTrain as an AI Quality Analyst evaluating a conversational AI personalization feature: design multi-turn prompts, rate side-by-side text responses, and write defensible rationales. 3-month contract, remote, 20+ hrs/week with required PST overlap.