Evaluate frontier AI coding agents by reviewing model-generated ETL pipelines, data warehouses, and distributed systems; contractor, 20+ hrs/week at $80/hr (USD). Ideal for data engineers with 2+ years' hands-on ETL and infrastructure experience.
Coding & Software
100% Remote Hourly · $80/hr
$80/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 29, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It’s where people start and grow careers teaching AI—discover open projects across the industry, build a profile, and apply in minutes. Creating an OpenTrain account is free.
About AI training and this work
AI training (data labeling and human evaluation) is the human side of building modern AI: people prepare, review, and judge examples that models learn from. In this role you will apply engineering judgment to assess model-generated code and infrastructure—directly shaping how coding agents and data platforms behave in production-like scenarios.
The role
OpenTrain AI is hiring a Data Engineer Coding Agent Evaluator to review model-generated data infrastructure and pipeline implementations. This is a part-time contractor role requiring 20+ hours per week, paid at $80/hour (USD). Work is remote and open worldwide; English fluency is required.
Employment type: Contractor, part-time
Time: 20+ hours per week
Rate: $80/hour (USD)
Location: Remote — worldwide (English required)
What you'll do
Evaluate frontier AI coding models on data engineering tasks, judging correctness and completeness.
Review model-generated ETL pipelines, data warehouse implementations, analytics components, and distributed data systems.
Identify bugs, edge cases, failure modes, and scalability or performance concerns in generated implementations.
Compare outputs from multiple frontier models and assess relative strengths and weaknesses.
Judge model-generated designs and code against real engineering standards and operational best practices.
Requirements
2+ years of professional data engineering experience.
Hands-on experience building ETL pipelines, data warehouses, analytics platforms, or distributed data systems.
Regular use or familiarity with AI coding agents such as Cursor, Claude Code, Codex, Windsurf, or Gemini CLI.
Ability to evaluate model-generated data infrastructure and pipeline implementations and spot bugs, edge cases, and scalability issues.
Experience operating large-scale data platforms is a plus.
Who should apply
This role is a fit for practicing data engineers, analytics engineers, platform engineers, or senior backend engineers who routinely design or review ETL and data platform code. You should be comfortable reading code, system designs, and deployment patterns and applying practical production standards when evaluating model outputs.
Data engineers who review pipeline code and schemas regularly.
Engineers experienced with distributed systems, storage, and analytics tooling.
Practitioners who have used AI coding assistants to accelerate development and can judge their outputs critically.
How it works
OpenTrain AI is the hiring and contracting organization for this role. Once you apply through OpenTrain, you’ll complete qualification tasks that test your ability to evaluate model-generated code and infrastructure. Work consists of short evaluation jobs where you read, compare, and rate outputs against engineering criteria. Payment is hourly as a contractor at the posted rate.
Data you’ll evaluate: model-generated programming and infrastructure code (computer code / programming).
Label type: evaluation and rating tasks that document correctness, risks, and improvement suggestions.
Application process: create an OpenTrain profile, complete qualifications, and accept tasks that fit your schedule.
Assess frontier AI coding agents by running realistic backend engineering tasks and reviewing model-generated code for correctness, maintainability, and performance. Remote contractor role paying $85/hr, 20+ hours/week, sprint-based work with typical tasks taking 2–3 hours.
Evaluate and improve agentic coding models by reviewing agent trajectories, verifying outputs, and designing rubrics on a 5-week remote contract. Requires 5+ years of hands-on software experience and daily overlap with PST.