Assess AI coding agents through realistic software tasks, trajectory reviews, testing, and technical judgments. This remote, two-month contractor role requires eight hours daily and four-hour PST overlap.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 31, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps people discover projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.
Work with cutting-edge AI systems
Build a portfolio of specialized evaluation experience
Find remote opportunities aligned with your technical background
About AI Training and Coding Evaluation
AI training is the human side of building artificial intelligence. People help modern models improve by preparing examples, reviewing outputs, testing behavior, and providing structured feedback. In this role, your software engineering judgment will help evaluate systems designed to perform development work.
Contribute to the improvement of generative AI coding systems
Apply practical engineering judgment to model-generated work
Work remotely in a fast-growing area of technology
The Role
OpenTrain is recruiting an Agentic Coding Model Evaluator to assess and improve AI systems that perform software development work. You will work with realistic coding tasks, inspect model trajectories, validate generated solutions, and provide technical judgments about correctness, reasoning, and behavior.
The work combines hands-on software engineering with structured model evaluation. It may also include designing benchmark tasks and creating task-specific evaluation rubrics.
Remote freelance contractor engagement
Two-month contract duration
Eight hours of work each day
Four-hour daily overlap with Pacific Standard Time
20+ hours per week listed as the time requirement
English-language work
What You’ll Do
You will execute coding tasks in agentic coding environments and determine whether the resulting implementations function as intended. The role requires close inspection of unfamiliar codebases, model behavior, test results, logs, and development artifacts.
Execute realistic coding tasks in agentic coding environments
Assess generated solutions for correctness, reasoning, and behavior
Review model trajectories and compare outputs to identify meaningful differences
Read unfamiliar codebases and run tests and terminal commands
Inspect logs and artifacts to validate implementation behavior
Write concise, evidence-based rationales for rankings and assessments
Follow detailed evaluation instructions, milestones, and quality standards
Design, test, and calibrate multi-step coding tasks
Develop evaluation rubrics tailored to specific tasks
Identify broken environments, ambiguous tasks, and evaluation issues
Escalate problems with supporting evidence
Required Qualifications
Although the opportunity is categorized as entry level in the supplied role data, the stated requirements call for at least five years of professional experience in software engineering, QA, developer tooling, data or ML engineering, or another code-intensive technical field.
Five or more years of professional experience in a code-intensive technical role
Strong practical proficiency in one or two programming languages or ecosystems
Ability to read and understand unfamiliar codebases
Strong debugging, testing, and edge-case reasoning skills
Ability to assess functional correctness consistently
Hands-on proficiency with Linux or Ubuntu and terminal workflows
Practical experience with Git, package managers, test runners, and related developer tooling
Familiarity with coding-agent workflows
Ability to evaluate AI-generated coding solutions consistently
Relevant Technical Ecosystems
Strong proficiency may come from one or two of the following programming languages or ecosystems. Your experience should support practical code review, debugging, testing, and evaluation rather than only theoretical familiarity.
Python
JavaScript or TypeScript
Rust
Java
C or C++
Bash
Haskell
Swift
SQL
Why Build an AI Training Career With OpenTrain
AI training and data-labeling work offers a way to contribute directly to how state-of-the-art AI systems behave. OpenTrain helps you build a durable professional profile so your evaluation experience can support future opportunities across this growing field.
Work remotely with a computer and internet connection
Build credible experience in AI model evaluation
Showcase specialized technical work in your OpenTrain profile
Discover projects that match your software and engineering skills
How to Apply
Create a free OpenTrain account, build your profile around your software engineering and evaluation experience, and apply in minutes. Be prepared to demonstrate practical proficiency with code, debugging, testing, developer tooling, and AI-generated coding assessments.
Apply through OpenTrain
Highlight relevant professional experience and programming ecosystems
Emphasize Linux, Git, testing, debugging, and code-review capabilities
Use your software engineering judgment to evaluate agentic coding models, verify solutions, and rank model trajectories. This flexible, worldwide contract role requires 20+ hours per week and strong hands-on programming experience.
Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.
Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.