Use your software engineering judgment to evaluate agentic coding models, verify solutions, and rank model trajectories. This flexible, worldwide contract role requires 20+ hours per week and strong hands-on programming experience.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects where human expertise helps improve the artificial intelligence systems people use every day.
Create a free OpenTrain account to build a professional AI training profile, discover projects that match your skills, and apply in minutes. Your work can become part of a durable portfolio in a fast-growing technology field.
About AI Training Work
AI training is the human side of building modern artificial intelligence. Contributors review examples, evaluate model outputs, write structured feedback, and test whether an AI system behaves correctly. This role focuses on the software engineering judgment needed to assess agentic coding models.
Work on cutting-edge systems that generate and execute software solutions
Apply practical technical expertise to model evaluation and dataset improvement
Contribute remotely with flexible project work designed for part-time schedules
The Role
OpenTrain is seeking an Agentic Coding Model Evaluation Specialist to evaluate and improve datasets used by agentic coding models. You will work with realistic software tasks in an agentic coding harness, review model trajectories, verify whether proposed solutions work, and produce precise annotations and assessments.
The role combines hands-on programming judgment with structured evaluation of model behavior. It is suited to experienced practitioners who can distinguish functional implementations from partial fixes, shallow solutions, and other incorrect outcomes.
Employment type: Contractor and part time
Time requirement: 20 or more hours per week
Language: English
Availability: Worldwide
Experience level listed for the role: Entry level
What You'll Do
You will execute realistic coding tasks while maintaining model blindness and session independence. You will follow task instructions, milestones, planned interactions, and evaluation guardrails consistently.
You will inspect unfamiliar code and validate proposed behavior using commands, tests, logs, generated artifacts, and targeted checks. Your assessments should be specific, evidence-based, and consistent across model trajectories.
You may also help create realistic multi-step coding tasks, define user intent and milestone structures, write task-specific rubrics and binary evaluation criteria, review completed work, and escalate broken environments or unclear instructions with supporting evidence.
Run commands and tests using command-line tools
Read unfamiliar codebases, scripts, logs, and generated artifacts
Debug issues and investigate edge cases
Verify whether proposed software solutions are functional
Rank and grade model-generated coding trajectories
Write clear rationales grounded in observable evidence
Design technical tasks that extend beyond tutorials, README flows, or simple bug fixes
Flag inconsistent instructions or broken environments with supporting evidence
Required Qualifications
This role requires substantial practical experience despite the listed entry-level classification. You should be comfortable making independent technical judgments about code behavior, implementation quality, and model-generated work.
Five or more years of experience in software engineering, QA, developer tooling, data engineering, ML engineering, or another code-heavy discipline
Strong hands-on experience with one or two production-relevant programming languages or ecosystems
Ability to understand unfamiliar codebases and interpret tests and scripts
Ability to use command-line tools, debug issues, and reason about edge cases
Judgment for assessing functional correctness and ranking model-generated coding trajectories
Clear technical writing and consistent evaluation judgment
Ability to design realistic multi-step coding tasks and task-specific evaluation rubrics
Helpful Background
The following experience is valuable for this work but is presented as helpful background rather than a required qualification.
AI training and data labeling are expanding ways to work in technology, including programming evaluation, model output review, and human feedback. Contributors help shape how advanced AI systems understand and produce software.
OpenTrain gives you a place to build a credible profile around this work, find projects aligned with your technical skills, and grow AI training experience into a longer-term career portfolio. Apply through OpenTrain to get started.
Remote work from anywhere with an internet connection
Flexible part-time work that can fit around other commitments
A free profile for showcasing relevant AI training experience
Opportunities to work directly on state-of-the-art AI systems
Assess AI coding agents through realistic software tasks, trajectory reviews, testing, and technical judgments. This remote, two-month contractor role requires eight hours daily and four-hour PST overlap.
Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.
Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.