Evaluate how an AI uses personal context to improve responses in Vietnamese. Create multi-turn prompts, rank model outputs, and write detailed rationales in a flexible remote contract.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this remote opportunity, helping contributors discover meaningful AI work and grow their experience in a fast-moving field.
Free to create an OpenTrain account
Remote contractor opportunity
Work focused on Vietnamese-language AI evaluation
About AI Training Work
AI training is the human side of building artificial intelligence. People review and evaluate model outputs so AI systems can become more accurate, relevant, and helpful. This work gives contributors a direct role in shaping how modern AI behaves while offering flexible opportunities that can fit around other commitments.
Contribute to the development of cutting-edge AI systems
Use language skills, analytical judgment, and attention to detail
Work remotely with a computer and reliable internet connection
The Role
OpenTrain is recruiting an AI Quality Analyst to evaluate a new personalization feature for Gemini. You will assess how effectively the model uses personal data from past conversations, Gmail, Search, and YouTube to produce more relevant and helpful responses.
This contractor role focuses on Vietnamese-language evaluation. You will create prompts from your own experiences, review personalized outputs, and assess them for Grounding, Integration, and Helpfulness.
Language focus: Vietnamese reading and writing
Experience level: Entry level
Pay: $15 per hour
Engagement: 3 months
Work arrangement: Fully remote
What You'll Do
You will test personalized AI behavior through realistic, creative conversations and provide structured judgments about the quality of each response. Your evaluations should reference specific conversation turns and clearly explain why one output is stronger than another.
Maintaining data hygiene is an important part of the role. Evaluation conversations must be deleted after use.
Design and execute multi-turn conversational prompts spanning 1 to 5 turns
Use personal information and experiences as context for prompts
Check responses for Grounding issues, including unsupported claims and hallucinations
Assess whether personal data is integrated naturally into the response
Rank two model responses side by side based on helpfulness and enjoyment
Write clear, structured rationales for your rankings
Delete evaluation conversations after use
Requirements
This role requires strong Vietnamese reading and writing ability, careful analysis of nuanced personalization, and the creativity to design prompts based on personal context. You should be comfortable making independent judgments about subtle differences in AI response quality and explaining those judgments in writing.
A desktop or laptop with reliable internet is required. The role is remote and requires flexibility to support global 24-hour operations.
High-level Vietnamese reading and writing proficiency
Experience evaluating AI model responses for personalization quality
Ability to create creative multi-turn prompts using personal context
Skill in side-by-side model response ranking and rationale writing
Strong analytical thinking about Grounding, Integration, and Helpfulness
Exceptional attention to detail
Excellent written communication
Self-motivation and ability to work independently
Desktop or laptop with reliable internet
Helpful Background
Experience in data annotation, AI quality evaluation, or content moderation can be helpful. A BS, BA, or equivalent experience in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related field is also helpful, but the listed experience level for this opportunity is entry level.
Data annotation experience
AI quality evaluation experience
Content moderation experience
Background in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related field
Schedule and Commitment
The engagement requires at least 4 hours per day and supports up to 40 hours per week. The listed time requirement is 20 or more hours per week, with 4 hours of overlap with Pacific Standard Time required.
Evaluate how well AI personalizes Thai-language conversations using creative multi-turn prompts, detailed ratings, and side-by-side response comparisons. This remote three-month contract pays $15 per hour.
Review how effectively AI uses personal context to produce relevant, grounded responses. This remote contractor role offers entry-level applicants flexible work of 20+ hours per week through OpenTrain.
Evaluate personalized AI responses in Indonesian by creating multi-turn prompts, ranking model outputs, and writing clear rationales. This remote contract pays $15 per hour and requires 20+ hours weekly.