Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It connects specialists with focused projects where they can apply professional expertise to emerging AI systems, build a profile, and grow a durable record of experience.
Creating an OpenTrain account is free, and candidates can apply in minutes. In this role, you will contribute specialized Python and software engineering judgment to the evaluation of AI coding agents.
About AI Coding Agent Evaluation
AI training is the human side of building artificial intelligence. Modern models learn from examples and reviews created by people, including technical evaluations of code, model outputs, tool calls, and software engineering workflows.
Coding-agent evaluation helps determine whether an AI system correctly understands and completes software engineering tasks. Your analysis can help identify errors, inefficient approaches, and opportunities to improve how these systems reason and act.
The Role
OpenTrain is recruiting a Python Coding Agent Evaluation Engineer to evaluate and improve next-generation AI coding agents. The work combines software engineering judgment, technical debugging, and AI model evaluation.
You will assess complete agentic coding trajectories, including prompts, tool calls, intermediate actions, code changes, execution results, and final outputs. This is a part-time contractor role requiring 20+ hours per week, with work available worldwide.
Role type: Part-time contractor
Time requirement: 20+ hours per week
Location: Worldwide and remote
Working language: English
Experience level: Entry level listing with substantial technical requirements
What You'll Do
You will review how coding agents interpret and execute software engineering tasks, then evaluate the technical quality of their approaches and results. You will use your engineering judgment to validate code and understand where an agent succeeds or breaks down.
AI coding agents such as Claude Code, Codex, Cursor, or equivalent tools may support your analysis, debugging, and validation work.
Review complete coding-agent trajectories from prompts through final outputs.
Evaluate technical correctness and identify errors or inefficient approaches.
Read, debug, test, and validate Python code produced or modified by AI agents.
Identify opportunities to improve AI model performance.
Work with LLM-powered agents, workflows, and development environments.
Requirements
This role requires strong hands-on Python expertise and professional software engineering experience. You should also have practical experience with AI or LLM engineering and regular exposure to AI coding agents.
At least five years of professional software engineering experience.
Strong hands-on expertise in Python software engineering.
Six months to one year of practical AI or LLM engineering experience.
Experience building agents, agent loops, LLM-powered applications, data pipelines, or similar systems.
Regular hands-on experience using AI coding agents.
Ability to evaluate tool use, intermediate actions, code changes, execution results, and final outputs.
Why Do This Work
AI training is a fast-growing way to work in tech, and contributors help shape how state-of-the-art AI systems behave. This project lets you apply your software engineering background directly to the development and evaluation of emerging coding-agent technology.
Remote, flexible AI training work can fit around other commitments while helping you build experience at the intersection of Python engineering, LLM systems, and model evaluation.
Apply professional Python expertise to cutting-edge AI systems.
Work remotely from anywhere in the world.
Contribute to the improvement of AI coding-agent behavior.
Build a focused AI training portfolio through OpenTrain.
How to Apply
Create a free OpenTrain account, build your profile, and apply in minutes. OpenTrain helps you organize specialized AI training experience into a portable professional portfolio as you work on projects in this growing field.
Evaluate Python coding agents by reviewing trajectories, tool use, code changes, execution results, and final outputs. This remote contractor role offers 20+ hours per week for experienced software engineers with AI or LLM expertise.
Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.
Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.