Skip to content
OpenTrain AIFor AI Companies

Python AI Coding Agent Evaluator

Evaluate next-generation AI coding agents by reviewing coding trajectories, debugging Python outputs, and validating technical results. This fully remote, one-month contract requires senior Python engineering experience and practical LLM or agent development expertise.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Expert

Experience

Aug 15, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover opportunities, build their professional profiles, and grow in a rapidly developing industry.

  • Create an OpenTrain account for free.
  • Work remotely on projects that help shape how advanced AI systems perform.

About AI Coding Agent Evaluation

AI training is the human side of building artificial intelligence. In coding-agent evaluation, experienced engineers inspect how AI systems interpret software tasks, use tools, modify code, and produce working results. Your technical judgment helps improve the reliability and capabilities of next-generation AI agents.

  • Contribute directly to the development of AI systems.
  • Apply professional software engineering expertise to cutting-edge model evaluation.
  • Work remotely in a fast-growing area of AI training and data labeling.

The Role

OpenTrain AI is seeking a Senior Python Engineer to evaluate and improve AI coding agents. You will review agentic coding trajectories, assess technical correctness, debug and validate Python outputs, and use LLM-powered coding tools as part of your workflow.

This is a fully remote, full-time contractual opportunity with a one-month contract duration. The schedule information calls for availability of 40 hours per week with required overlap with the global team; the engagement is also listed as requiring 20+ hours per week.

  • Role: Python AI Coding Agent Evaluator
  • Experience level: Expert
  • Work arrangement: Fully remote
  • Engagement: Contractual
  • Duration: 1 month
  • Language: English
  • Availability: 20+ hours per week, with 40 hours per week specified for the full-time schedule

What You'll Do

You will analyze complete agent workflows rather than reviewing code in isolation. This includes examining prompts, tool calls, intermediate actions, code changes, execution results, and the agent's overall approach to software engineering tasks.

  • Review agentic coding trajectories, including prompts, tool calls, intermediate actions, code changes, and execution results.
  • Evaluate whether AI agents correctly understand and execute software engineering tasks.
  • Identify technical errors and opportunities to improve agent behavior.
  • Read, debug, test, and validate Python code produced or modified by AI agents.
  • Work with LLM-powered agents, workflows, and development environments.
  • Use Claude Code, Codex, Cursor, or similar coding agents to support analysis and validation.

Required Qualifications

This role is designed for an experienced software engineer with strong Python skills and practical experience developing or working with AI and LLM-powered systems. You should be comfortable assessing both implementation quality and the intermediate decisions made by an AI coding agent.

  • 5+ years of professional software engineering experience.
  • Strong hands-on expertise in Python.
  • At least 6 months to 1 year of practical AI or LLM engineering experience.
  • Experience building agents, agent loops, or LLM-powered applications.
  • Hands-on experience with AI coding agents such as Claude Code, Codex, Cursor, or equivalent tools.
  • Ability to evaluate agentic workflows, including tool usage, intermediate execution steps, code modifications, and resulting outputs.

Helpful Background

The strongest candidates will combine advanced Python programming ability with a working understanding of how LLM-powered agents operate. Familiarity with coding-agent workflows will help you quickly analyze technical behavior and validate results.

  • Strong Python programming skills.
  • Familiarity with LLM-powered agents and agentic workflows.
  • Experience using Claude Code, Codex, Cursor, or similar coding agents.

Remote Contract Schedule

This opportunity is fully remote and contractual. You will need to maintain the required weekly availability and overlap with the global team during the engagement.

  • Fully remote work.
  • Full-time contractual opportunity.
  • One-month contract.
  • 40 hours per week specified, with 20+ hours per week listed as the minimum time requirement.
  • Required overlap with the global team.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Agentic Coding Annotator, Model Evaluation

Evaluate and improve agentic coding models by reviewing agent trajectories, verifying outputs, and designing rubrics on a 5-week remote contract. Requires 5+ years of hands-on software experience and daily overlap with PST.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 29, 2026

Python Engineer, AI Coding-Tool Evaluation (Part-Time)

Part-time contractor role testing and evaluating an internal AI coding tool; $100/hr, remote, under 20 hrs/week. Ideal for intermediate Python engineers with hands-on experience using AI coding assistants like Cursor, Windsurf, or Claude Code.

Coding & Software
Computer Code Programming
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $100/hr

Posted Oct 17, 2025

Python AI Model Training & Evaluation Engineer

Build and evaluate AI systems through Python development, supervised fine-tuning, RLHF, response ranking, and public-data analysis. This fully remote contractor role offers 20, 30, or 40 hours per week for a one-month engagement.

Coding & Software
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026