Evaluate how AI personalizes responses for Thai users using multi-turn prompts, personal context, and side-by-side rankings. This remote contract pays $15/hour and requires 20+ hours weekly.
Generative AI & RLHF
100% Remote Hourly · $15/hr
$15/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this remote AI training role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes. Creating an OpenTrain account is free.
Remote contract work available worldwide
Part-time schedule of 20+ hours per week
Hourly pay of $15 USD
About AI Training Work
AI training is the human side of building modern artificial intelligence. Contributors evaluate model responses, write prompts, review language, and provide structured feedback that helps AI systems become more accurate, useful, and natural.
Work directly with cutting-edge AI evaluation tasks
Use analytical judgment to assess model quality
Build experience in a fast-growing technology field
The Role
OpenTrain is seeking an intermediate AI Quality Analyst to evaluate a new AI personalization feature. You will assess whether the model uses information from past conversations, email, search, and video activity appropriately to make responses more relevant and helpful.
This remote, part-time contract focuses on Thai-language evaluation. You will need to use your primary personal Google account and enable personal data sources so the feature can be assessed using genuine personal context.
Role: AI Quality Analyst
Focus: Thai AI personalization evaluation
Employment type: Part-time contractor
Workload: 20+ hours per week
Pay: $15 USD per hour
Experience level: Intermediate
What You’ll Do
You will design realistic evaluation scenarios, assess model behavior across multiple turns, and document clear evidence for your judgments. The work requires careful attention to personalization quality, data use, naturalness, and overall usefulness.
Design and execute multi-turn conversational prompts spanning one to five turns.
Create starting prompts that require the AI to use personal information and experiences.
Evaluate responses against the intent expressed in the starting prompt.
Analyze grounding issues, integration quality, and overall helpfulness.
Stack-rank two model responses side by side based on helpfulness, ease of use, and enjoyment.
Write clear, defensible comparison rationales that reference specific turn numbers.
Extract and verify debug information to confirm that chat summaries and data sources were used properly.
Delete evaluation conversations to maintain strict data hygiene and prevent contamination of future chat history.
Required Skills
This role requires strong Thai reading and writing ability, thoughtful analysis of nuanced AI behavior, and the judgment to distinguish genuinely useful personalization from incorrect, forced, or poorly supported connections.
Read and write Thai with a high degree of competence.
Evaluate ambiguous AI responses for personalization quality.
Design creative multi-turn prompts based on personal context.
Identify incorrect personalization, poor inferences, and forced connections.
Spot subtle differences in naturalness, overnarration, and response quality.
Write clear, concise, and structured rationales for model rankings.
Use a primary personal Google account and enable personal data sources.
Have a desktop or laptop with a reliable internet connection.
Helpful Background
A BS or BA degree, or equivalent experience, in a relevant analytical field is helpful. Relevant backgrounds may include policy, law, ethics, linguistics, journalism, computer science, or a related field.
Experience in data annotation, AI quality evaluation, content moderation, or a related role is also helpful.
BS or BA degree or equivalent experience in an analytical field
Experience evaluating AI responses for personalization quality
Background in data annotation, AI quality evaluation, or content moderation
Experience with creative prompt engineering using personal context
Why This Work Matters
Every major AI system depends on people who prepare examples, evaluate outputs, and identify where models fall short. In this role, your Thai-language judgments will help assess how naturally and responsibly AI uses personal context to improve its responses.
Contribute to the development of more relevant AI experiences
Apply Thai-language expertise to advanced model evaluation
Work remotely with a flexible part-time structure
How to Apply
Create a free OpenTrain account, build your profile, and apply in minutes. Be ready to highlight your Thai proficiency, analytical experience, prompt design skills, and ability to write evidence-based evaluation rationales.
Review the role requirements before applying.
Emphasize Thai reading and writing proficiency.
Describe relevant AI evaluation, annotation, moderation, or analytical experience.
Confirm that you have a desktop or laptop and reliable internet access.
Evaluate personalized AI responses in Turkish, create multi-turn prompts, and provide clear, rubric-based feedback. This remote contractor role pays $15 per hour and offers 30 or 40 hours per week.
Use your Thai-English communication, research, and analytical skills to help improve large language models. Analyze content, validate claims, answer questions, and create detailed training feedback remotely for 20+ hours per week.
Evaluate how AI personalizes conversations in Hindi by designing multi-turn prompts, ranking responses, and writing evidence-based rationales. This remote contractor role pays $15 per hour and requires 20+ hours weekly.