Evaluate how an AI personalization feature uses information from conversations and digital activity to create helpful German responses. Work remotely as a contractor for $15 per hour, with 20+ hours available weekly.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 15, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free. This contractor opportunity lets you contribute directly to the development and evaluation of modern AI systems from a remote setting.
Remote contractor engagement
Part-time schedule with 20+ hours per week
Three-month engagement
Global team operating across a 24-hour schedule
About AI Training and Quality Evaluation
AI training is the human side of building artificial intelligence. Contributors write prompts, review model responses, and provide structured judgments that help AI systems become more accurate, relevant, and useful.
In this role, your evaluations will focus on personalization: whether a model uses relevant information appropriately while avoiding unsupported assumptions. Your careful comparisons and written reasoning will help shape how AI responds to people in real-world situations.
Work on a cutting-edge AI personalization feature
Use creativity and analytical judgment together
Help improve the quality and usefulness of AI-generated responses
Work remotely with a flexible part-time structure
The Role
OpenTrain is recruiting an AI Quality Analyst to evaluate a new personalization feature for Gemini. You will assess how effectively the model uses personal information from past conversations, Gmail, Google Search, and YouTube activity to deliver more relevant and helpful responses.
The work combines creative prompt design with rigorous quality analysis. You will create prompts based on your own experiences, inspect how the model applies personalization, and explain your conclusions clearly in German and English as applicable to the evaluation workflow.
Role focus: AI personalization quality evaluation
Focus language: German
Experience level: Entry level
Data type: Text
Evaluation format: Response rating and side-by-side comparison
What You'll Do
You will evaluate multi-turn conversations, typically lasting one to five turns, and determine whether the model understood and used relevant personal context. Your reviews should be specific, evidence-based, and attentive to both the user's intent and the quality of the final experience.
Design and execute multi-turn conversational prompts that require the AI to use personal information and experiences.
Evaluate model responses against your intended prompt and determine whether personalization was applied appropriately.
Analyze grounding issues and check whether claims about you are supported by evidence rather than flawed inferences or hallucinations.
Assess whether personal data is integrated naturally without robotic or excessive narration.
Stack-rank two model responses side by side based on overall helpfulness, ease of use, and enjoyment.
Write clear, defensible rationales that reference specific issues or strengths in the conversation.
Maintain strict data hygiene by deleting evaluation conversations.
Requirements
This role requires high-level German reading and writing ability, strong analytical judgment, and the discipline to explain nuanced decisions in a structured way. You should be comfortable working independently in a remote environment and using a desktop or laptop with a reliable internet connection.
High proficiency in reading and writing German
Experience designing creative prompts for AI evaluation
Ability to evaluate responses for grounding and integration quality
Exceptional analytical thinking when reviewing nuanced AI outputs
Strong evaluation judgment for personalization quality
Meticulous attention to detail during side-by-side comparisons
Excellent written communication for clear, structured rationales
Self-motivation and ability to work independently
Desktop or laptop with a good internet connection
Helpful Background
A bachelor's degree or equivalent experience in a relevant field may be helpful, but the role emphasizes practical evaluation ability, language competence, and careful reasoning. Previous experience with AI quality work can help you contribute effectively.
BS or BA degree, or equivalent experience, in Linguistics, Computer Science, Policy, Journalism, or a related field
Experience in data annotation, AI quality evaluation, or content moderation
Compensation, Schedule, and Application
This is a contractor engagement paying $15 per hour. The engagement is scheduled for three months, with a commitment of at least four hours per day and up to 40 hours per week. You must be available for four hours of overlap with Pacific Standard Time.
Apply through OpenTrain to take the next step toward flexible, remote work helping improve how advanced AI systems understand context and respond to people.
$15 per hour
At least 4 hours per day
Up to 40 hours per week
At least 20 hours per week based on the stated time requirement
Review how effectively AI uses personal context to produce relevant, grounded responses. This remote contractor role offers entry-level applicants flexible work of 20+ hours per week through OpenTrain.
Evaluate how well an AI personalization feature uses conversational and account context to produce relevant Japanese responses. Work remotely as an independent contractor for $15 per hour with a 20+ hour weekly commitment.
Evaluate how well an AI model uses Russian-language personal context, compare responses, and write clear rationales. Work remotely for 20+ hours per week at $15 per hour.