Skip to content
OpenTrain AIFor AI Companies

Agentic Coding Model Evaluator

Assess AI coding agents through realistic software tasks, trajectory reviews, testing, and technical judgments. This remote, two-month contractor role requires eight hours daily and four-hour PST overlap.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 31, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps people discover projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.

  • Work with cutting-edge AI systems
  • Build a portfolio of specialized evaluation experience
  • Find remote opportunities aligned with your technical background

About AI Training and Coding Evaluation

AI training is the human side of building artificial intelligence. People help modern models improve by preparing examples, reviewing outputs, testing behavior, and providing structured feedback. In this role, your software engineering judgment will help evaluate systems designed to perform development work.

  • Contribute to the improvement of generative AI coding systems
  • Apply practical engineering judgment to model-generated work
  • Work remotely in a fast-growing area of technology

The Role

OpenTrain is recruiting an Agentic Coding Model Evaluator to assess and improve AI systems that perform software development work. You will work with realistic coding tasks, inspect model trajectories, validate generated solutions, and provide technical judgments about correctness, reasoning, and behavior.

The work combines hands-on software engineering with structured model evaluation. It may also include designing benchmark tasks and creating task-specific evaluation rubrics.

  • Remote freelance contractor engagement
  • Two-month contract duration
  • Eight hours of work each day
  • Four-hour daily overlap with Pacific Standard Time
  • 20+ hours per week listed as the time requirement
  • English-language work

What You’ll Do

You will execute coding tasks in agentic coding environments and determine whether the resulting implementations function as intended. The role requires close inspection of unfamiliar codebases, model behavior, test results, logs, and development artifacts.

  • Execute realistic coding tasks in agentic coding environments
  • Assess generated solutions for correctness, reasoning, and behavior
  • Review model trajectories and compare outputs to identify meaningful differences
  • Read unfamiliar codebases and run tests and terminal commands
  • Inspect logs and artifacts to validate implementation behavior
  • Write concise, evidence-based rationales for rankings and assessments
  • Follow detailed evaluation instructions, milestones, and quality standards
  • Design, test, and calibrate multi-step coding tasks
  • Develop evaluation rubrics tailored to specific tasks
  • Identify broken environments, ambiguous tasks, and evaluation issues
  • Escalate problems with supporting evidence

Required Qualifications

Although the opportunity is categorized as entry level in the supplied role data, the stated requirements call for at least five years of professional experience in software engineering, QA, developer tooling, data or ML engineering, or another code-intensive technical field.

  • Five or more years of professional experience in a code-intensive technical role
  • Strong practical proficiency in one or two programming languages or ecosystems
  • Ability to read and understand unfamiliar codebases
  • Strong debugging, testing, and edge-case reasoning skills
  • Ability to assess functional correctness consistently
  • Hands-on proficiency with Linux or Ubuntu and terminal workflows
  • Practical experience with Git, package managers, test runners, and related developer tooling
  • Familiarity with coding-agent workflows
  • Ability to evaluate AI-generated coding solutions consistently

Relevant Technical Ecosystems

Strong proficiency may come from one or two of the following programming languages or ecosystems. Your experience should support practical code review, debugging, testing, and evaluation rather than only theoretical familiarity.

  • Python
  • JavaScript or TypeScript
  • Rust
  • Java
  • C or C++
  • Bash
  • Haskell
  • Swift
  • SQL

Why Build an AI Training Career With OpenTrain

AI training and data-labeling work offers a way to contribute directly to how state-of-the-art AI systems behave. OpenTrain helps you build a durable professional profile so your evaluation experience can support future opportunities across this growing field.

  • Work remotely with a computer and internet connection
  • Build credible experience in AI model evaluation
  • Showcase specialized technical work in your OpenTrain profile
  • Discover projects that match your software and engineering skills

How to Apply

Create a free OpenTrain account, build your profile around your software engineering and evaluation experience, and apply in minutes. Be prepared to demonstrate practical proficiency with code, debugging, testing, developer tooling, and AI-generated coding assessments.

  • Apply through OpenTrain
  • Highlight relevant professional experience and programming ecosystems
  • Emphasize Linux, Git, testing, debugging, and code-review capabilities

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Agentic Coding Model Evaluation Specialist

Use your software engineering judgment to evaluate agentic coding models, verify solutions, and rank model trajectories. This flexible, worldwide contract role requires 20+ hours per week and strong hands-on programming experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026

Python Coding Agent Evaluation Engineer

Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026

Python AI Coding Agent Evaluator

Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.

Coding & Software
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Aug 15, 2026