Review how effectively AI uses personal context to produce relevant, grounded responses. This remote contractor role offers entry-level applicants flexible work of 20+ hours per week through OpenTrain.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Sep 4, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the platform where people find and build careers in AI training and data labeling, create a professional profile, and apply to projects that help shape modern AI systems.
Remote contract opportunity
Part-time engagement with a 20+ hour weekly commitment
Worldwide availability
English-language work
About AI Training and Quality Evaluation
AI training is the human work behind better artificial intelligence. Contributors write prompts, review model outputs, compare responses, and provide structured feedback so AI systems become more accurate, useful, and natural.
Help improve how AI understands and uses context
Evaluate nuanced model behavior rather than simply checking right or wrong answers
Work remotely with flexible scheduling
The Role
As an AI Personalization Quality Analyst, you will evaluate a personalization feature by testing how well an AI model uses information from a person's previous conversations, email, search activity, and video activity to make responses more relevant and helpful.
You will create multi-turn prompts from your own experiences, assess personalized responses across Grounding, Integration, and Helpfulness, and compare model outputs side by side. This entry-level contractor role requires at least 20 hours per week.
Employment type: Contractor and part-time
Experience level: Entry level
Work location: Worldwide and remote
Focus language: English
Data type: Text
What You’ll Do
Your evaluations will focus on whether personalization is accurate, relevant, natural, and genuinely useful. You will document your reasoning clearly so model quality can be assessed consistently.
Design and execute multi-turn conversational prompts, typically one to five turns, using personal context
Evaluate responses against the intent of the starting prompt
Check whether claims about the user are grounded in supporting evidence
Identify flawed inferences, hallucinations, incorrect personalization, and forced connections
Assess whether personal information is integrated naturally without excessive or robotic narration
Rank two model responses side by side based on helpfulness, ease of use, and overall quality
Write clear, concise, and defensible rationales referencing specific conversation turns
Delete evaluation conversations to maintain strict data hygiene and prevent them from affecting future chat history
Requirements
This role is designed for careful, analytical contributors who can recognize subtle differences in AI-generated language and explain their judgments in writing. English proficiency is required because English is the focus language.
High degree of English reading and writing proficiency
Exceptional analytical thinking when evaluating nuanced or ambiguous personalized responses
Creative prompt engineering ability for multi-turn prompts based on personal context
Strong evaluation judgment and understanding of personalization quality
Ability to identify poor inferences, incorrect personalization, and forced connections
Meticulous attention to detail when comparing responses side by side
Excellent written communication and structured reasoning
Ability to work independently in a remote setting
Helpful Background
A bachelor’s degree or equivalent experience in a relevant analytical field is helpful but not required. Previous experience in AI evaluation or annotation can also strengthen your application.
Policy, law, ethics, linguistics, journalism, computer science, or a related analytical background
Experience with data annotation
Experience with AI quality evaluation
Experience in content moderation or a related role
How to Apply
Create a free OpenTrain account to build your AI training profile and apply through OpenTrain. This opportunity is available to candidates worldwide who meet the English-language and weekly time requirements.
Review the role requirements
Highlight analytical, writing, prompt design, or evaluation experience
Evaluate how naturally AI systems use personal context in Polish conversations. Create multi-turn prompts, rank responses, and write detailed feedback in a remote, one-month contractor project paying $20 per hour.
Evaluate how well an AI personalization feature uses conversational and account context to produce relevant Japanese responses. Work remotely as an independent contractor for $15 per hour with a 20+ hour weekly commitment.
Evaluate personalized AI responses in Indonesian by creating multi-turn prompts, ranking model outputs, and writing clear rationales. This remote contract pays $15 per hour and requires 20+ hours weekly.