Evaluate and improve agentic coding models by reviewing agent trajectories, verifying outputs, and designing rubrics on a 5-week remote contract. Requires 5+ years of hands-on software experience and daily overlap with PST.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 29, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized AI-training projects, consolidate work across opportunities, and build a lasting freelance portfolio focused on data labeling, model evaluation, and human feedback.
OpenTrain AI is the hiring and contracting organization for this role. Creating an OpenTrain account is free and gives you access to projects like this one — flexible, remote, and directly shaping how advanced AI systems behave.
About AI training work in coding
AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. For coding-focused projects, contributors interact with model-generated code, rank trajectories, run and debug code, and create evaluation tasks that teach models to reason and act correctly.
This work is typically remote and flexible, and it gives software practitioners a hands-on way to influence frontier model capabilities in reasoning, tooling, and code generation.
The role
As an Agentic Coding Annotator you will evaluate and improve datasets used to train agentic coding models. You will execute and verify realistic coding tasks within an agentic harness, compare and rank blinded model trajectories, and produce high-quality annotations and rubrics for offline evaluations.
The role blends live, online evaluations (interacting with blinded models and ranking trajectories) and offline work (designing multi-step tasks, creating rubrics, and grading generated outputs). Your work directly advances model coding and reasoning capabilities.
What you'll do
Execute realistic coding tasks inside the assigned agentic coding harness while maintaining model blindness and session independence.
Verify model outputs by reading code, running commands, checking logs, and inspecting generated artifacts.
Perform targeted validation of outputs using tests, scripts, and manual checks.
Write clear, evidence-based rationales for trajectory rankings and assessments.
Design multi-step coding tasks for offline evaluations, including user intent, milestones, and success criteria.
Create and refine task-specific rubrics and binary evaluation criteria.
Review completed work for quality, completeness, consistency, and schema compliance.
Identify and escalate broken environments or unclear instructions with supporting evidence.
Requirements
You must meet the core experience and technical requirements below. These are strict: they ensure you can run and judge realistic engineering work and produce defensible annotations.
5+ years of experience in software engineering, QA, developer tooling, data/ML engineering, or a similar code-heavy role.
Strong hands-on experience in at least one or two production-relevant languages (examples: Python, JavaScript/TypeScript, Rust, Java, C/C++, Bash/CLI, Haskell, Swift, SQL).
Ability to read, understand, and debug unfamiliar codebases.
Ability to run and interpret tests, scripts, and CLI tools.
Skill in debugging, reasoning about edge cases, and assessing whether an implementation is functionally correct.
Helpful background
The following are not required but will make you a stronger candidate and allow you to contribute faster and at higher quality.
Strong Docker skills and experience building or debugging reproducible environments.
Experience working in large, complex repositories (beyond small or greenfield projects).
Demonstrated originality and sound engineering judgment when defining technical problems.
Experience designing realistic, multi-step coding tasks that go beyond tutorials or trivial bug fixes.
Contract details & schedule
Contract duration: 5 weeks.
Core schedule: 8 hours per day with a required 4-hour overlap with Pacific Time (PST). The project lists a minimum time expectation of 20+ hours/week; candidates should expect the stated daily schedule and be able to meet overlap requirements.
Employment type: Contractor, part-time. This role is open worldwide; English proficiency is required. Compensation is not specified in this posting.
Who should apply & next steps
Apply if you are a practiced software engineer who enjoys reading and debugging code, writing clear technical judgments, and designing evaluation tasks that test models’ real engineering abilities.
To be considered, prepare examples or a brief description of relevant experience (large/repo work, Docker, testing/debugging). OpenTrain will manage contracting and onboarding for successful candidates.
Join OpenTrain to train next-generation coding assistants by running, evaluating, and documenting agent-driven software workflows. Part-time remote contract work (20+ hrs/week) paying $70–$126/hr for hands-on technical testing, code review, and process documentation.
Review AI-generated Angular code, assess correctness and best practices, and provide clear written feedback to improve model outputs. Part-time contractor role, $25/hr, remote worldwide, under 20 hours/week.