Evaluate how AI uses personal context in Polish-language conversations, compare model responses, and write clear quality rationales. This remote contractor role pays $20 per hour for about 4 hours daily.
Generative AI & RLHF
100% Remote Hourly · $20/hr
$20/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is recruiting contractors for projects where human judgment helps improve the quality, usefulness, and reliability of modern AI systems.
Apply through OpenTrain and build experience in a fast-growing AI training industry.
Work remotely with flexible project opportunities designed for contributors around the world.
About AI Evaluation Work
AI evaluation is the human side of building artificial intelligence. Contributors write prompts, review model responses, compare outputs, and explain which answers are more helpful or accurate so AI systems can improve.
Help assess how AI handles real conversational context.
Use careful reasoning to identify strong responses, weak inferences, and hallucinations.
Contribute to cutting-edge model development through structured human feedback.
The Role
OpenTrain is recruiting an AI Personalization Evaluation Analyst for a project focused on evaluating a Gemini personalization feature. You will create multi-turn prompts based on your own experiences, assess how the model uses personal signals from conversations and connected services, compare paired responses, and write defensible rationales focused on grounding, integration, and helpfulness.
Contractor engagement
Remote work available worldwide
Polish fluency required for reading and writing evaluation content
Entry-level role with relevant analytical or evaluation experience welcomed
$20 per hour
About 4 hours per day and up to 40 hours per week
One-month engagement
Availability in your local time zone with 4 hours of overlap with PST
What You'll Do
You will evaluate personalization behavior through short, realistic conversations and document your decisions clearly. Strong performance requires attention to detail, consistent judgment, and disciplined organization of evaluation data.
Create short multi-turn prompts, typically one to five turns, that test personalization behavior.
Judge whether the model uses personal context appropriately.
Identify flawed inferences and hallucinations.
Compare two responses side by side and select the more helpful, natural, and usable answer.
Write concise explanations that reference specific conversation turns.
Explain quality differences using grounding, integration, and helpfulness criteria.
Check supporting debug information.
Maintain strict data hygiene so future evaluation conversations are not affected.
Requirements
This role requires fluency in Polish and the ability to make nuanced judgments about AI response quality. You should be comfortable working independently with personal-context tasks, structured evaluation criteria, and written rationales.
Fluency in Polish for reading and writing evaluation notes.
Experience judging AI response quality using grounding and helpfulness criteria.
Ability to design multi-turn prompts from personal context.
Experience with prompt design, personalization review, or side-by-side response evaluation.
Comfort comparing model outputs and writing clear, defensible rationales.
Independence, clear communication, and reliable remote work habits.
Helpful Background
Prior experience in AI training is helpful but not limited to one career path. Background in data annotation, AI quality evaluation, or content moderation may be relevant, as may a BS, BA, or equivalent experience in an analytical field.
Data annotation experience
AI quality evaluation experience
Content moderation experience
Analytical background in policy, law, ethics, linguistics, journalism, or computer science
Equivalent practical experience in a relevant analytical field
How to Apply
Create a free OpenTrain account, build your profile, and apply in minutes. This project offers a focused way to gain hands-on experience evaluating how AI systems use conversational context while working remotely on a defined contractor engagement.
Confirm Polish fluency and the required evaluation experience.
Review the one-month, $20-per-hour contractor engagement.
Apply through OpenTrain with your relevant background and availability.
Review how effectively AI uses personal context to produce relevant, grounded responses. This remote contractor role offers entry-level applicants flexible work of 20+ hours per week through OpenTrain.
Evaluate how AI personalizes conversations in Hindi by designing multi-turn prompts, ranking responses, and writing evidence-based rationales. This remote contractor role pays $15 per hour and requires 20+ hours weekly.
Evaluate how well an AI personalization feature uses conversational and account context to produce relevant Japanese responses. Work remotely as an independent contractor for $15 per hour with a 20+ hour weekly commitment.