Help evaluate next-generation AI coding agents by reviewing trajectories, tool calls, code changes, and execution results. This remote, one-month contract requires strong Python expertise, five years of software engineering experience, and 20+ hours per week.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 30, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain is the hiring and contracting organization for this opportunity, helping specialists apply their expertise to projects that shape how advanced AI systems work.
Create a free OpenTrain account and apply in minutes.
Build a profile that reflects your AI training and technical experience.
Develop a durable portfolio through cutting-edge AI work.
About AI Coding Agent Evaluation
AI training is the human side of building artificial intelligence. For coding agents, experienced engineers review prompts, actions, generated code, tool use, and outcomes so AI systems can become more accurate, efficient, and dependable in software development environments.
Review how AI agents interpret and execute software engineering tasks.
Assess code, workflows, intermediate actions, and final outputs.
Contribute directly to the improvement of LLM-powered development tools.
The Role
OpenTrain is recruiting a Python Coding Agent Evaluation Engineer to evaluate and improve next-generation AI coding agents. You will combine software engineering judgment with hands-on review of agentic coding trajectories, including prompts, tool calls, code modifications, execution results, and final responses.
This is a fully remote, full-time contractual opportunity structured as a focused one-month project. You will collaborate with a global team and are expected to contribute 20+ hours per week.
Role: Python Coding Agent Evaluation Engineer
Engagement: One-month contractual project
Work arrangement: Fully remote
Time requirement: 20+ hours per week
Language: English
Location: Worldwide
What You'll Do
You will determine whether AI agents correctly understand and execute software engineering tasks. Your evaluations will cover the technical quality of each trajectory, the efficiency of the agent's approach, the correctness of code changes, and whether execution results support the final outcome.
Review complete coding-agent trajectories from prompts through final responses.
Inspect tool calls, intermediate actions, code modifications, and execution steps.
Read, debug, test, and validate Python code generated or modified by AI agents.
Identify technical errors, inefficient approaches, and opportunities to improve agent behavior.
Evaluate how agents use tools and manage intermediate steps.
Assess technically sound outcomes from Claude Code, Codex, Cursor, or equivalent coding agents.
Requirements
This role requires substantial professional software engineering experience and practical exposure to AI or LLM engineering. Candidates should be comfortable evaluating both the implementation details and overall behavior of coding agents.
At least five years of professional software engineering experience.
Six months to one year of practical AI or LLM engineering experience.
Experience building agents, agent loops, LLM-powered applications, data pipelines, or similar systems.
Regular hands-on use of AI coding agents such as Claude Code, Codex, Cursor, or equivalent tools.
Ability to evaluate agentic workflows, tool calls, code changes, and execution results.
Skill in debugging, testing, and validating Python code produced by AI agents.
Who Should Apply
This opportunity is suited to software engineers who enjoy analyzing how complex systems behave and can make precise technical judgments about AI-generated code. It may be a strong fit if you have built or regularly used LLM-powered development tools and want to help shape the next generation of coding agents.
Experienced Python engineers with strong debugging and testing skills.
Developers familiar with agent loops, coding-agent workflows, or LLM applications.
Engineers who can assess both individual code changes and complete execution trajectories.
Professionals available for a focused remote project requiring 20+ hours weekly.
How It Works
Create a free OpenTrain account, complete your profile, and apply for this contract in minutes. OpenTrain helps people discover and build careers in AI training and data labeling, a fast-growing field where human expertise improves modern AI systems.
If selected, you will work remotely with a global team during the focused one-month engagement. Your work will contribute to evaluating and improving AI coding agents while strengthening your technical AI training portfolio.
Create or update your OpenTrain profile.
Highlight your Python, software engineering, and AI or LLM experience.
Apply through OpenTrain for consideration.
Complete the remote one-month contract with a 20+ hour weekly commitment.
Evaluate Python coding agents by reviewing trajectories, tool use, code changes, execution results, and final outputs. This remote contractor role offers 20+ hours per week for experienced software engineers with AI or LLM expertise.
Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.
Create and evaluate realistic software-engineering benchmarks for coding agents using production-like repositories, secure coding, and rigorous testing. This fully remote contractor assignment runs 4 to 8 weeks.