Evaluate how conversational AI uses personal context in Russian, compare responses, and write evidence-based feedback. This remote, three-month contractor role pays $15 per hour and requires at least 20 hours weekly.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 24, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is recruiting for this contractor engagement and gives contributors a place to discover opportunities, build a professional AI training profile, and apply in minutes.
Creating an OpenTrain account is free. Your work in this role can help you build credible experience in AI response evaluation and grow a portfolio in a rapidly developing field.
About AI Training Work
AI training is the human side of building artificial intelligence. People evaluate model responses, write prompts, annotate examples, and provide feedback that helps AI systems become more accurate, useful, and natural.
This remote work can offer flexible part-time opportunities while allowing contributors to work directly on the quality of modern AI systems.
The Role
OpenTrain AI is seeking a Russian AI Personalization Evaluator to assess personalized, multi-turn conversational AI interactions. You will test how models use personal context, determine whether their responses are helpful and natural, and provide clear feedback supported by evidence.
The evaluation work focuses on identifying inaccurate personalization, weak or unsupported inferences, hallucinations, forced connections, unnecessary narration, and other quality issues.
Remote contractor engagement lasting 3 months
$15 per hour
Part-time schedule requiring at least 20 hours per week
Availability of at least 4 hours per day, up to 40 hours per week
At least 4 hours of overlap with Pacific Time required
What You'll Do
You will create and assess realistic conversations designed to test whether an AI system understands and applies personal context appropriately. Your judgments should be precise, independent, and grounded in specific evidence from the conversation and available debugging information.
Design creative, multi-turn prompts grounded in personal context.
Evaluate whether responses follow the intent of the starting prompt.
Determine whether personalization is accurate, relevant, useful, and natural.
Check whether claims about the user are supported by available evidence.
Identify hallucinations, unsupported inferences, forced connections, and unnecessary narration.
Compare and rank responses side by side for helpfulness, ease of use, naturalness, and overall quality.
Write concise, defensible rationales that reference specific conversation turns.
Review debugging information to verify that conversation summaries and data sources were used correctly.
Delete evaluation conversations to maintain careful data hygiene.
Requirements
This role requires strong Russian reading and writing ability, careful analytical judgment, and the ability to explain subtle differences between AI responses. A reliable desktop or laptop and internet connection are also required.
The evaluation process requires using a primary personal Google account and enabling relevant personal data sources for the evaluation.
Proficient written and reading ability in Russian
Ability to create creative, multi-turn prompts from personal context
Strong reasoning for nuanced AI response judgments
Ability to recognize inaccurate personalization, hallucinations, unsupported inferences, and flawed connections
Excellent written communication and attention to detail
Ability to provide constructive feedback and detailed annotations independently
Reliable desktop or laptop and internet connection
Willingness to use a primary personal Google account and relevant personal data sources
Helpful Background
A bachelor's degree or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or a related analytical field is helpful but not required. Previous experience in data annotation, AI quality evaluation, content moderation, or a related area is also valuable.
How to Apply
OpenTrain helps people start and grow careers in AI training and data labeling. Create a free OpenTrain account, build your profile, and apply for this Russian AI Personalization Evaluator opportunity in minutes.
Review the role requirements and schedule.
Highlight your Russian proficiency and analytical evaluation experience.
Apply through OpenTrain and build your AI training portfolio.
Use your Russian fluency and written judgment to evaluate sensitive AI model behavior, identify adversarial patterns, and improve model safety. This remote contractor role offers 7 hours per week at $38-$42 per hour.
Evaluate Polish-language AI conversations, compare personalized responses, and write detailed quality rationales in a remote one-month contract paying $20 per hour.
Review how conversational AI uses personal context in German, compare responses, and explain nuanced quality judgments. This remote contractor role offers $15 per hour for a three-month engagement.