Assess frontier AI coding agents by running realistic backend engineering tasks and reviewing model-generated code for correctness, maintainability, and performance. Remote contractor role paying $85/hr, 20+ hours/week, sprint-based work with typical tasks taking 2–3 hours.
Coding & Software
100% Remote Hourly · $85/hr
$85/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 29, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people find and build careers in AI training and data labeling. We help freelancers discover projects, consolidate opportunities, and build a unified portfolio so they can grow a durable freelance career in the fast-moving AI training industry.
About AI training work
AI training (also called data labeling or human feedback work) is the human side of building AI: people prepare, review, and judge examples that teach modern models how to behave. These roles are often remote, flexible, and accessible to people with professional or domain experience rather than formal ML credentials.
Work is typically 100% remote and flexible, suitable for part-time schedules.
Contributors shape how state-of-the-art systems perform in real engineering scenarios.
The role
OpenTrain is hiring a Backend Engineering Code Agent Evaluator to assess frontier AI coding models by completing realistic backend engineering tasks with AI coding agents and reviewing their outputs. This is evaluation and judgment work—you will not be building production features but applying professional software engineering standards to model-generated code.
This is a contractor, part-time role requiring 20+ hours per week. Work is sprint-based with 12–24 hour stretches depending on project needs. Typical tasks take about 2–3 hours after an initial ramp-up. Compensation is paid per accepted task at $85 USD per hour.
Employment type: Contractor, Part-time
Time requirement: 20+ hours/week
Pay: $85 USD per hour, paid for accepted work
Work format: Remote, sprint-based engagements with short focused stretches
What you'll do
Use frontier AI coding agents to complete realistic backend engineering tasks and exercises.
Review model-generated code for correctness, quality, maintainability, and performance.
Identify bugs, edge cases, architectural tradeoffs, and failure modes in model outputs.
Compare outputs from multiple models and summarize strengths and weaknesses.
Apply professional backend engineering judgment to evaluate solutions rather than produce production-ready features.
Requirements
2+ years of professional backend engineering experience.
Hands-on experience building APIs, distributed systems, microservices, backend platforms, or databases.
Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, or Gemini CLI.
Strong ability to evaluate model-generated code and spot bugs, edge cases, and architectural tradeoffs.
Experience with large-scale production systems is preferred.
Fluent English required for task instructions and reports.
Who should apply
This role is ideal for backend engineers who enjoy critical code review, can reason about system design and failure modes, and are comfortable working with AI-assisted coding tools. If you prefer applying engineering judgment over building features day-to-day, this evaluation role is a strong fit.
Engineers who regularly use AI code assistants and can assess their outputs critically.
People seeking flexible, remote part-time contracting work in the AI training industry.
How it works
You will receive sprint assignments that describe backend tasks and evaluation criteria. After a short ramp-up, typical assignments take 2–3 hours. Submit your evaluations and findings; accepted work is paid at the stated hourly rate.
OpenTrain manages hiring, payment, and project matching. Work is worldwide and remote; applicants must be able to work in English and commit to the stated time requirements.
Assignments include clear acceptance criteria; payment is tied to accepted deliverables.
Expect short, focused sprints (12–24 hour windows) and flexible scheduling within the weekly hour commitment.
Evaluate frontier AI coding agents by reviewing model-generated ETL pipelines, data warehouses, and distributed systems; contractor, 20+ hrs/week at $80/hr (USD). Ideal for data engineers with 2+ years' hands-on ETL and infrastructure experience.
Join OpenTrain as a contractor reviewing model-generated ML code, MLOps, and LLM application workflows—apply engineering judgment to find bugs, edge cases, and performance issues. Part-time remote work at $85/hr, 20+ hours/week, open worldwide to experienced ML engineers.