Evaluate Python coding agents by reviewing trajectories, tool use, code changes, execution results, and final outputs. This remote contractor role offers 20+ hours per week for experienced software engineers with AI or LLM expertise.
Coding & Software
100% Remote
Worldwide
Eligibility
Intermediate
Experience
Aug 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help professionals discover specialized projects, build a focused portfolio, and grow practical experience in a rapidly developing field. Creating an OpenTrain account is free, and this opportunity is available remotely worldwide.
Remote contractor opportunity
Part-time engagement with 20+ hours per week
Open to candidates worldwide
English-language work
About AI Coding Agent Evaluation
AI training is the human side of building artificial intelligence. For coding agents, experienced engineers review model behavior, code, tool use, and execution outcomes to help determine whether an agent can complete software engineering tasks accurately and efficiently.
This work combines hands-on programming with structured evaluation of LLM-powered workflows and development environments. Your technical judgment can help identify successful strategies, failure points, and practical opportunities to improve how AI agents operate.
Review real AI-generated software engineering behavior
Assess technical correctness and efficiency
Contribute practical feedback to AI development
Work remotely in a growing AI training field
The Role
OpenTrain is recruiting a Python AI Coding Agent Evaluation Engineer to support the development and evaluation of AI agents. You will examine complete agentic coding trajectories and determine whether an agent understands and executes software engineering tasks correctly.
The role involves reviewing prompts, tool calls, intermediate actions, code modifications, execution results, and final outputs. You will use software engineering judgment to distinguish technically correct solutions from flawed or inefficient approaches.
Role level: Intermediate
Work type: Contractor and part-time
Time requirement: 20+ hours per week
Work location: Worldwide and fully remote
What You'll Do
You will evaluate how AI coding agents reason through and complete development tasks. Reviews will cover the full trajectory from the initial prompt through tool usage, code changes, execution, debugging, and the final result.
The work requires reading, debugging, testing, and validating Python code produced or modified by AI agents. You will also assess LLM-powered workflows and development environments through structured technical review.
Review prompts, tool calls, intermediate actions, and final outputs
Assess whether agents correctly understand and execute software engineering tasks
Identify technical errors, inefficient methods, and improvement opportunities
Read, debug, test, and validate Python code
Evaluate code modifications, execution results, and technical outputs
Use Claude Code, Codex, Cursor, or comparable AI coding tools
Analyze how agent actions affect the final result
Required Qualifications
This opportunity is intended for experienced software engineers with strong Python development skills and practical experience working with AI or LLM-based systems. You should be comfortable interpreting agent behavior, tool usage, intermediate execution steps, code changes, and resulting outputs.
At least 5 years of professional software engineering experience
Experience building agents, agent loops, LLM-powered applications, data pipelines, or similar systems
Regular experience using AI coding agents in software development workflows
Ability to evaluate technical correctness, efficiency, and outcomes
Helpful Background
Experience building agentic systems, LLM-powered applications, data pipelines, or related AI engineering systems is particularly relevant. Familiarity with multiple coding agents is useful because the evaluation work may cover behavior across different development environments.
Strong candidates can explain why an agent's approach succeeds or fails and how its intermediate actions influence the final output.
Experience with Claude Code, Codex, Cursor, or equivalent tools
Background building agentic systems or agent loops
Experience developing LLM-powered applications
Experience working with data pipelines
Ability to explain flawed, correct, or inefficient technical approaches
Why Work With OpenTrain
AI training and data-labeling work gives technical professionals a direct role in shaping how modern AI systems behave. OpenTrain connects you with specialized evaluation projects, supports practical portfolio development, and helps you build a focused freelance career in AI development.
You will work remotely with professionals across a global network while contributing software engineering judgment to cutting-edge AI systems.
Build practical experience evaluating advanced AI systems
Contribute to the development of coding agents and LLM workflows
Work remotely with a global professional network
Build a durable portfolio of AI development contributions
Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.
Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.
Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.