Skip to content
OpenTrain AIFor AI Companies

Agentic Coding Annotator, Model Evaluation

Evaluate and improve agentic coding models by reviewing agent trajectories, verifying outputs, and designing rubrics on a 5-week remote contract. Requires 5+ years of hands-on software experience and daily overlap with PST.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Jul 29, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized AI-training projects, consolidate work across opportunities, and build a lasting freelance portfolio focused on data labeling, model evaluation, and human feedback.

OpenTrain AI is the hiring and contracting organization for this role. Creating an OpenTrain account is free and gives you access to projects like this one — flexible, remote, and directly shaping how advanced AI systems behave.

About AI training work in coding

AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. For coding-focused projects, contributors interact with model-generated code, rank trajectories, run and debug code, and create evaluation tasks that teach models to reason and act correctly.

This work is typically remote and flexible, and it gives software practitioners a hands-on way to influence frontier model capabilities in reasoning, tooling, and code generation.

The role

As an Agentic Coding Annotator you will evaluate and improve datasets used to train agentic coding models. You will execute and verify realistic coding tasks within an agentic harness, compare and rank blinded model trajectories, and produce high-quality annotations and rubrics for offline evaluations.

The role blends live, online evaluations (interacting with blinded models and ranking trajectories) and offline work (designing multi-step tasks, creating rubrics, and grading generated outputs). Your work directly advances model coding and reasoning capabilities.

What you'll do

  • Execute realistic coding tasks inside the assigned agentic coding harness while maintaining model blindness and session independence.
  • Verify model outputs by reading code, running commands, checking logs, and inspecting generated artifacts.
  • Perform targeted validation of outputs using tests, scripts, and manual checks.
  • Write clear, evidence-based rationales for trajectory rankings and assessments.
  • Design multi-step coding tasks for offline evaluations, including user intent, milestones, and success criteria.
  • Create and refine task-specific rubrics and binary evaluation criteria.
  • Review completed work for quality, completeness, consistency, and schema compliance.
  • Identify and escalate broken environments or unclear instructions with supporting evidence.

Requirements

You must meet the core experience and technical requirements below. These are strict: they ensure you can run and judge realistic engineering work and produce defensible annotations.

  • 5+ years of experience in software engineering, QA, developer tooling, data/ML engineering, or a similar code-heavy role.
  • Strong hands-on experience in at least one or two production-relevant languages (examples: Python, JavaScript/TypeScript, Rust, Java, C/C++, Bash/CLI, Haskell, Swift, SQL).
  • Ability to read, understand, and debug unfamiliar codebases.
  • Ability to run and interpret tests, scripts, and CLI tools.
  • Skill in debugging, reasoning about edge cases, and assessing whether an implementation is functionally correct.

Helpful background

The following are not required but will make you a stronger candidate and allow you to contribute faster and at higher quality.

  • Strong Docker skills and experience building or debugging reproducible environments.
  • Experience working in large, complex repositories (beyond small or greenfield projects).
  • Demonstrated originality and sound engineering judgment when defining technical problems.
  • Experience designing realistic, multi-step coding tasks that go beyond tutorials or trivial bug fixes.

Contract details & schedule

Contract duration: 5 weeks.

Core schedule: 8 hours per day with a required 4-hour overlap with Pacific Time (PST). The project lists a minimum time expectation of 20+ hours/week; candidates should expect the stated daily schedule and be able to meet overlap requirements.

Employment type: Contractor, part-time. This role is open worldwide; English proficiency is required. Compensation is not specified in this posting.

Who should apply & next steps

Apply if you are a practiced software engineer who enjoys reading and debugging code, writing clear technical judgments, and designing evaluation tasks that test models’ real engineering abilities.

To be considered, prepare examples or a brief description of relevant experience (large/repo work, Docker, testing/debugging). OpenTrain will manage contracting and onboarding for successful candidates.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Agentic AI Coding Expert (Remote, Part-Time)

Join OpenTrain to train next-generation coding assistants by running, evaluating, and documenting agent-driven software workflows. Part-time remote contract work (20+ hrs/week) paying $70–$126/hr for hands-on technical testing, code review, and process documentation.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $70–$126/hr

Posted Jul 9, 2026

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

Angular Expert, AI Code Evaluation & Review

Review AI-generated Angular code, assess correctness and best practices, and provide clear written feedback to improve model outputs. Part-time contractor role, $25/hr, remote worldwide, under 20 hours/week.

Coding & Software
Computer Code Programming
Remote · Worldwide
Part-time · Flexible
Entry level
Hourly · $25/hr

Posted Mar 7, 2025