Evaluate Spanish conversational AI responses for accurate, natural personalization, comparing multi-turn outputs and writing clear rationales. This remote contractor role offers $15 per hour and requires 20+ hours weekly.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. As an OpenTrain contractor, you can discover projects, build a profile that reflects your experience, and apply in minutes through one free account.
- Work with OpenTrain AI as a part-time contractor.
- Build experience in a fast-growing AI training industry.
- Manage your AI training career from a single profile.
About AI Training Work
AI systems improve through human feedback and careful evaluation. Contributors review model outputs, test how well systems follow instructions, and explain what makes an answer accurate, useful, natural, and safe. This work directly influences how modern conversational AI behaves.
- Remote work using a computer and internet connection.
- Flexible project work that can fit around other commitments.
- Hands-on experience evaluating cutting-edge AI systems.
The Role
OpenTrain is recruiting a Spanish Personalized AI Response Evaluator to assess personalized conversational AI interactions. You will create multi-turn prompts grounded in personal experiences, evaluate how naturally and accurately a model uses relevant information, and compare responses for grounding, integration, helpfulness, and overall usability.
This intermediate-level contractor role pays $15 per hour and requires 20 or more hours per week. The work is worldwide and part time, with Spanish written proficiency as a core requirement. It also requires willingness to use a primary personal Google account and enable relevant personal data sources for evaluation.
- Role: Spanish Personalized AI Response Evaluator
- Work type: Part-time contractor
- Pay: $15 USD per hour
- Time commitment: 20+ hours per week
- Location: Worldwide, remote
- Data type: Text
- Focus: RLHF and evaluation rating
What You'll Do
You will assess whether personalized conversational AI responses use personal context appropriately and remain faithful to the user's intent. Evaluations may involve conversations of one to five turns and require both creative prompt design and careful quality judgment.
You will also verify debugging information to confirm that summaries and relevant data sources were used correctly. After review, you will maintain data hygiene by deleting evaluation conversations.
- Design and execute conversational prompts using personal context.
- Create multi-turn prompts, typically spanning one to five turns.
- Check whether responses follow the intent of each prompt.
- Identify unsupported claims, hallucinations, flawed inferences, and forced connections.
- Spot unnatural over-explanation and incorrect personalization.
- Compare responses side by side for grounding, integration, helpfulness, ease of use, and enjoyment.
- Write concise, defensible rationales tied to specific conversation turns.
- Provide constructive feedback and verify summaries and relevant data sources.
- Delete evaluation conversations after completing reviews.
Requirements
Strong Spanish reading and writing ability is required. You should be able to analyze nuanced or ambiguous AI responses, create creative multi-turn prompts from personal context, and recognize when personalization is incorrect, unsupported, or forced.
Clear, structured written communication and close attention to detail are essential. You must be able to work independently in a remote setting with a reliable computer and internet connection.
- Strong written Spanish proficiency.
- Excellent analytical thinking for nuanced and ambiguous responses.
- Ability to design creative multi-turn prompts from personal context.
- Ability to identify unsupported personalization and flawed inferences.
- Ability to compare responses for grounding, integration, and helpfulness.
- Ability to write clear, turn-specific evaluation rationales.
- Reliable computer and internet connection.
- Willingness to use a primary personal Google account and enable relevant personal data sources.
Helpful Background
A bachelor's degree or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or a related analytical field is helpful. Previous experience in data annotation, AI quality evaluation, content moderation, or related work is also valuable.
- Policy, law, ethics, linguistics, journalism, or computer science background.
- Experience with data annotation or AI quality evaluation.
- Experience in content moderation or related analytical work.
- Intermediate experience level.
How to Apply
Create a free OpenTrain account and apply through the platform in minutes. Your profile can help showcase relevant analytical, language, and AI evaluation experience as you build a longer-term portfolio in AI training and data labeling.
- Apply for this worldwide remote contractor opportunity through OpenTrain.
- Highlight Spanish writing, evaluation, prompt design, and analytical experience.
- Use the role to develop credible experience in conversational AI evaluation.