Evaluate AI chatbot responses using realistic business prompts across Google Workspace. Help improve personalization, accuracy, relevance, and usefulness in a flexible, part-time remote contract role for US-based business owners.
Generative AI & RLHF
Remote
1 country
Eligibility
Entry
Experience
Aug 4, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the leading platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and grow in a fast-moving technology field.
Free account creation
Opportunities to build experience in AI training and evaluation
A direct way to contribute to the development of modern AI systems
About AI Training Work
AI training is the human side of building artificial intelligence. Evaluators review model outputs, test AI systems with realistic prompts, and provide structured feedback so models become more accurate, relevant, and useful.
Work directly with cutting-edge AI tools
Use human judgment to assess model quality
Help shape how AI systems respond to real-world needs
The Role
OpenTrain AI is seeking a Workspace Personalization AI Evaluator to assess AI-generated responses using personalized prompts across Google Workspace applications. You will create business-related scenarios, interact with AI chatbots, and evaluate whether their responses are accurate, relevant, complete, personalized, and useful.
This role focuses on Workspace Personalization Evals and is designed for a US-based business owner with more than 10 employees who has substantial experience with small business operations and Google Workspace.
Contractor and part-time engagement
20+ hours per week
English-language work
US-based applicants only
Entry-level experience designation
What You’ll Do
You will follow structured evaluation guidelines while testing AI chatbot conversations and documenting your findings. Your feedback will support model improvements by identifying where AI responses succeed or fail in realistic business contexts.
Create realistic business-related prompts based on defined user goals
Interact with multiple AI chatbots in conversations of up to five turns
Assess response clarity, usefulness, accuracy, relevance, personalization, and completeness
Provide structured feedback and comparative evaluations
Identify incorrect context retrieval, missing information, and irrelevant responses
Submit conversation transcripts and evaluation results
Requirements
This role requires practical business experience and strong familiarity with Google Workspace. You should be comfortable using AI tools, applying objective judgment, and following detailed evaluation instructions while maintaining confidentiality.
Business owner with more than 10 employees
Strong understanding of small business operations
Strong analytical and critical thinking skills
Gmail must be your primary business communication tool
High familiarity with Gmail, Chat, Calendar, Drive, Docs, Sheets, Slides, and Meet
Heavy usage of Google Chat, Drive, and Calendar
Comfort interacting with AI tools and chatbots
Ability to follow structured evaluation guidelines
Ability to provide clear, objective feedback
Ability to maintain confidentiality
Why This Work Matters
Every major AI system depends on people who can assess outputs against real-world expectations. By bringing your small-business experience to personalized Workspace evaluations, you will help improve how AI understands context and supports business users.
Apply real operational judgment to AI evaluation
Work remotely with flexible part-time potential
Gain experience in a rapidly growing AI training industry
Evaluate a personalization feature in Polish by designing short multi-turn prompts, comparing paired model responses, and writing concise, defensible quality rationales. Contractor role, $20/hr, remote, ~4 hours/day with 4-hour overlap with PST for a 1-month engagement.
Join OpenTrain to evaluate a new personalization feature in Japanese: design multi-turn prompts, compare paired model responses, and write clear rationales in a remote, 3-month contractor role paying $15/hr with PST overlap.
Evaluate personalized AI responses for relevance, accuracy, helpfulness, and appropriate use of context. This remote, part-time contractor role is open to U.S.-based applicants who actively use email, calendar, photo, and file-management applications.