Evaluate ChatGPT and Claude through realistic business workflows, score their outputs, and provide actionable feedback. This remote contractor role pays $30-$90 per hour and requires 20+ hours per week.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free. Your profile can help you showcase relevant AI training experience and grow a career in a fast-moving field where human judgment directly improves how AI systems work.
About AI Training Work
AI training is the human side of building artificial intelligence. People evaluate model responses, write feedback, and judge whether AI outputs are accurate, useful, complete, and relevant. This work helps shape the behavior and reliability of modern AI systems.
This opportunity focuses on evaluating generative AI assistants in realistic professional settings. It is remote and offers flexible contractor work for contributors who can commit at least 20 hours per week.
The Role
As an AI Agent Workflow Evaluator, you will test AI assistants such as ChatGPT and Claude through complex, multi-step business workflows. You will assess how well each assistant handles practical operational requirements and document the results clearly.
Your evaluations will help identify strengths, weaknesses, and opportunities to improve the quality, completeness, relevance, and reliability of next-generation AI systems. Previous AI training experience is not required.
- Employment type: Remote contractor and part-time
- Time commitment: 20+ hours per week
- Compensation: $30-$90 per hour
- Work authorization: United States, Canada, United Kingdom, Ireland, Australia, or New Zealand
- United States strongly preferred
- Working language: English
What You'll Do
You will reproduce authentic workplace use cases by connecting AI assistants with business and productivity tools. You will maintain accurate records so that your evaluations are transparent, consistent, and reproducible.
Using defined rubrics and evaluation criteria, you will score AI-generated outputs and provide constructive feedback that points to specific improvements. You will also track recurring patterns in assistant behavior across different workflows.
- Run complex, multi-step business scenarios that reflect genuine professional workflows
- Use ChatGPT, Claude, or both to complete realistic tasks
- Record each step of a workflow and compare AI behavior with operational requirements
- Score generated outputs using rubrics, QA scorecards, grading criteria, or review guidelines
- Write precise, actionable feedback about model performance
- Identify recurring strengths, weaknesses, and process opportunities
- Connect AI assistants with tools such as Google Drive, Gmail, Slack, and Notion
- Maintain detailed written documentation that supports transparency and reproducibility
Requirements
This role is suited to professionals who understand how technology supports real business operations and who can assess work against clear quality standards. You should be comfortable investigating multi-step processes, making careful judgments, and explaining your reasoning in written English.
- At least five years of professional experience in a business function where technology is used to solve operational challenges
- Completed bachelor's degree or higher in any discipline
- Daily, hands-on professional use of ChatGPT, Claude, or both
- Experience creating, applying, or reviewing evaluation rubrics, QA scorecards, grading criteria, or content review guidelines
- Comfort connecting AI assistants with workplace and productivity software
- Excellent written English
- Ability to document findings precisely and provide clear, actionable feedback
Helpful Background
Experience in any of the following areas may help you contribute effectively. These backgrounds can provide useful practice in applying consistent standards, reviewing outputs, or documenting complex processes.
- Academic grading
- Quality assurance
- Hiring scorecards
- Content moderation
- Annotation guidelines
- AI model evaluation
- Documenting multi-step professional processes
- Judging outputs for quality, completeness, and relevance
How to Apply
Create or update your free OpenTrain profile and apply through OpenTrain AI. Highlight your professional experience, hands-on use of ChatGPT or Claude, evaluation or quality-review work, and ability to document complex workflows.
If selected, you will work as a remote contractor evaluating AI assistants in practical business scenarios. This flexible opportunity lets you contribute directly to the development of increasingly capable AI systems.