Build Python backends that replicate real SaaS tools or create realistic tasks and rubrics for evaluating AI agents. This remote, four-week contractor engagement pays $300 per approved task and requires at least one approved task daily.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. As the hiring and contracting organization for this role, OpenTrain connects skilled contributors with practical work that helps shape how advanced AI systems operate.
This engagement is designed for contractors who want to apply backend engineering judgment to frontier AI agent development while building credible experience in a fast-growing technical field.
About AI Training and Agent Evaluation
AI training is the human side of building artificial intelligence. Engineers and evaluators create examples, workflows, tests, and feedback that help AI systems become more capable, reliable, and useful.
In this role, your software engineering work supports AI agents by reproducing real tool behavior, creating realistic long-horizon tasks, and defining clear standards for judging whether an agent completed work correctly.
The Role
OpenTrain is recruiting contractors for one of two related tracks: building faithful Python connectors that replicate SaaS tools, or developing realistic long-horizon tasks and evaluation rubrics for AI agent performance.
You will work remotely for a four-week engagement starting immediately. The role requires a minimum of 20 hours per week and at least one completed and approved task each day.
- Engagement type: Remote contractor and part-time
- Duration: Four weeks
- Start: Immediately
- Time requirement: 20+ hours per week
- Compensation: $300 per approved task
- Additional approved tasks may be completed during the engagement
- Work location: Worldwide
- Working language: English
What You’ll Do
You may contribute to connector development, AI-agent task creation and evaluation, or both. The work combines hands-on Python backend engineering with careful validation of workflows, edge cases, and quality standards.
- Build Python backend applications that reproduce the behavior of tools such as Slack, Linear, Jira, Notion, Gmail, and wikis.
- Extend existing connectors and develop new connectors that faithfully emulate their target systems.
- Test implementations and validate connector behavior against expected workflows.
- Mine workflows and data to support realistic long-horizon task development.
- Author AI-agent tasks and verify their realism and correctness.
- Conduct structured quality assurance and participate in calibration and quality reviews.
- Write rubrics defining correct, partially correct, and deficient work.
- Use AI coding agents throughout development, QA, and validation.
- Meet agreed quality standards for implementation, testing, and evaluation.
Requirements
This role is listed at the entry level, but the source requirements call for at least three years of backend development experience. Candidates should be comfortable delivering production-minded Python services and using AI coding assistants as part of their regular engineering workflow.
- At least three years of backend development experience
- Strong Python skills with FastAPI, Flask, or Django
- Experience building scalable backends, REST APIs, and microservices
- Practical proficiency with Git, Docker, pipelines, testing, and code-quality practices
- Daily experience using AI coding assistants such as Claude Code, Codex, Cursor, GitHub Copilot, or similar tools
- Ability to validate SaaS connector behavior
- Ability to create realistic AI-agent tasks with clear evaluation rubrics
Helpful Background
Experience translating real software workflows into reliable APIs, test cases, or evaluation criteria will be useful across both work tracks. Strong judgment about correctness, realism, edge cases, and quality standards will help you contribute effectively to connector validation and AI-agent task evaluation.
- Workflow analysis and software behavior replication
- API design, backend testing, or service integration
- Quality assurance and structured evaluation
- Writing precise acceptance criteria or grading rubrics
- Comfort reviewing edge cases and ambiguous outcomes
Why Build an AI Training Career with OpenTrain
AI training and data-labeling work spans programming, evaluation, writing, search, language, images, audio, and more. Many projects are remote and flexible, while technical projects like this one allow experienced contributors to apply specialized skills directly to cutting-edge AI development.
OpenTrain helps contributors find and build careers in this growing industry. You can create a profile, discover relevant projects, apply in minutes, and develop a portfolio of experience over time. Creating an OpenTrain account is free.