Skip to content
OpenTrain AIFor AI Companies

Agentic Coding Model Evaluation Specialist

Use your software engineering judgment to evaluate agentic coding models, verify solutions, and rank model trajectories. This flexible, worldwide contract role requires 20+ hours per week and strong hands-on programming experience.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects where human expertise helps improve the artificial intelligence systems people use every day.

Create a free OpenTrain account to build a professional AI training profile, discover projects that match your skills, and apply in minutes. Your work can become part of a durable portfolio in a fast-growing technology field.

About AI Training Work

AI training is the human side of building modern artificial intelligence. Contributors review examples, evaluate model outputs, write structured feedback, and test whether an AI system behaves correctly. This role focuses on the software engineering judgment needed to assess agentic coding models.

  • Work on cutting-edge systems that generate and execute software solutions
  • Apply practical technical expertise to model evaluation and dataset improvement
  • Contribute remotely with flexible project work designed for part-time schedules

The Role

OpenTrain is seeking an Agentic Coding Model Evaluation Specialist to evaluate and improve datasets used by agentic coding models. You will work with realistic software tasks in an agentic coding harness, review model trajectories, verify whether proposed solutions work, and produce precise annotations and assessments.

The role combines hands-on programming judgment with structured evaluation of model behavior. It is suited to experienced practitioners who can distinguish functional implementations from partial fixes, shallow solutions, and other incorrect outcomes.

  • Employment type: Contractor and part time
  • Time requirement: 20 or more hours per week
  • Language: English
  • Availability: Worldwide
  • Experience level listed for the role: Entry level

What You'll Do

You will execute realistic coding tasks while maintaining model blindness and session independence. You will follow task instructions, milestones, planned interactions, and evaluation guardrails consistently.

You will inspect unfamiliar code and validate proposed behavior using commands, tests, logs, generated artifacts, and targeted checks. Your assessments should be specific, evidence-based, and consistent across model trajectories.

You may also help create realistic multi-step coding tasks, define user intent and milestone structures, write task-specific rubrics and binary evaluation criteria, review completed work, and escalate broken environments or unclear instructions with supporting evidence.

  • Run commands and tests using command-line tools
  • Read unfamiliar codebases, scripts, logs, and generated artifacts
  • Debug issues and investigate edge cases
  • Verify whether proposed software solutions are functional
  • Rank and grade model-generated coding trajectories
  • Write clear rationales grounded in observable evidence
  • Design technical tasks that extend beyond tutorials, README flows, or simple bug fixes
  • Flag inconsistent instructions or broken environments with supporting evidence

Required Qualifications

This role requires substantial practical experience despite the listed entry-level classification. You should be comfortable making independent technical judgments about code behavior, implementation quality, and model-generated work.

  • Five or more years of experience in software engineering, QA, developer tooling, data engineering, ML engineering, or another code-heavy discipline
  • Strong hands-on experience with one or two production-relevant programming languages or ecosystems
  • Ability to understand unfamiliar codebases and interpret tests and scripts
  • Ability to use command-line tools, debug issues, and reason about edge cases
  • Judgment for assessing functional correctness and ranking model-generated coding trajectories
  • Clear technical writing and consistent evaluation judgment
  • Ability to design realistic multi-step coding tasks and task-specific evaluation rubrics

Helpful Background

The following experience is valuable for this work but is presented as helpful background rather than a required qualification.

  • Docker experience
  • Building or debugging reproducible environments
  • Working in large, complex repositories
  • Defining non-trivial technical problems
  • Designing realistic tasks beyond tutorials, README flows, or simple bug fixes

Why Build an AI Training Career With OpenTrain

AI training and data labeling are expanding ways to work in technology, including programming evaluation, model output review, and human feedback. Contributors help shape how advanced AI systems understand and produce software.

OpenTrain gives you a place to build a credible profile around this work, find projects aligned with your technical skills, and grow AI training experience into a longer-term career portfolio. Apply through OpenTrain to get started.

  • Remote work from anywhere with an internet connection
  • Flexible part-time work that can fit around other commitments
  • A free profile for showcasing relevant AI training experience
  • Opportunities to work directly on state-of-the-art AI systems

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Agentic Coding Model Evaluator

Assess AI coding agents through realistic software tasks, trajectory reviews, testing, and technical judgments. This remote, two-month contractor role requires eight hours daily and four-hour PST overlap.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 31, 2026

Python Coding Agent Evaluation Engineer

Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026

Python Engineer, AI Coding Agent Evaluation

Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026