Skip to content
OpenTrain AIFor AI Companies

Python Engineer, AI Coding Agent Evaluation

Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 30, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects that help improve modern artificial intelligence systems.

Through OpenTrain, you can build a profile, showcase relevant experience, discover projects that match your skills, and apply in minutes. Creating an OpenTrain account is free.

  • Worldwide opportunity
  • Part-time contractor engagement
  • 20+ hours per week
  • English-language work

About AI Coding Agent Evaluation

AI training is the human side of building artificial intelligence. In coding-agent evaluation, experienced engineers examine how AI systems interpret software tasks, use tools, modify code, and respond to execution results.

Your technical judgment helps identify whether an AI coding agent reached a correct outcome, made sound software engineering decisions, and followed an efficient path. This work contributes to the development of next-generation AI systems while offering flexible, remote opportunities in a fast-growing field.

  • Review real AI-generated software engineering work
  • Assess technical correctness and agent behavior
  • Help improve the quality of AI coding systems
  • Work remotely with flexible part-time hours

The Role

OpenTrain is recruiting a Python Engineer, AI Coding Agent Evaluation contributor to assess complete agentic coding trajectories from the initial prompt through the final output. You will examine prompts, tool calls, intermediate actions, code changes, execution results, and final responses.

The role is listed as entry level, while the required profile includes at least five years of professional software engineering experience and six months to one year of practical AI or LLM engineering experience.

  • Contractor and part-time position
  • 20+ hours per week
  • Worldwide eligibility
  • Core focus: Python and AI coding-agent evaluation

What You’ll Do

You will analyze coding-agent workflows in detail, combining software engineering judgment with practical validation. The work involves reading, debugging, testing, and validating Python code produced or modified by AI agents.

  • Review coding-agent trajectories from the initial prompt through the final output
  • Evaluate technical correctness and the quality of software engineering decisions
  • Analyze tool usage, intermediate execution steps, code modifications, and resulting behavior
  • Read, debug, test, and validate Python code produced or modified by AI agents
  • Identify technical errors, inefficient approaches, and opportunities to improve agent performance
  • Use coding agents such as Claude Code, Codex, Cursor, or equivalent tools during analysis and validation
  • Assess whether agentic workflows and technical outcomes are correct

Requirements

This role requires strong hands-on Python expertise and the ability to explain whether an AI coding agent's technical outcome is correct. You should be comfortable evaluating both the final code and the sequence of decisions that produced it.

  • At least five years of professional software engineering experience
  • Strong hands-on Python software engineering expertise
  • Six months to one year of practical AI or LLM engineering experience
  • Experience building agents, agent loops, LLM-powered applications, data pipelines, or similar systems
  • Regular experience using AI coding agents in software development workflows
  • Ability to evaluate tool calls, intermediate execution steps, code modifications, and technical outputs
  • Fluent English communication for this English-language project

Who Should Apply

This opportunity is suited to Python engineers who want to apply their software development and AI engineering experience to the evaluation of emerging coding agents. It may be a strong fit if you routinely use tools such as Claude Code, Codex, Cursor, or equivalent systems and can distinguish correct, efficient engineering work from flawed or incomplete approaches.

The project is open worldwide and designed for contributors available to work at least 20 hours per week in a part-time contractor capacity.

  • Experienced Python software engineers
  • Engineers who have built or evaluated AI and LLM-powered systems
  • Developers who regularly use AI coding agents
  • Contributors who can investigate technical details and explain their conclusions

Build Your AI Training Career With OpenTrain

AI training and data-labeling work spans coding, language, audio, images, search, and model feedback. Contributors help prepare and review the examples that modern AI systems learn from, making this a direct way to participate in how advanced technology is built.

Apply through OpenTrain to begin building a durable AI training portfolio. Your profile can help you show credible experience, find relevant opportunities, and grow your career in this rapidly expanding industry.

  • Create a free OpenTrain account
  • Build a profile around your Python and AI engineering experience
  • Apply to this opportunity in minutes
  • Develop a portfolio of specialized AI training work

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Python AI Coding Agent Evaluation Engineer

Evaluate Python coding agents by reviewing trajectories, tool use, code changes, execution results, and final outputs. This remote contractor role offers 20+ hours per week for experienced software engineers with AI or LLM expertise.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Intermediate level

Posted Aug 30, 2026

Python Coding Agent Evaluation Engineer

Use your Python engineering expertise to evaluate next-generation coding agents, review complete trajectories, debug generated code, and assess tool use and execution results. Work remotely as a part-time contractor for 20+ hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026

Python AI Coding Agent Evaluator

Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.

Coding & Software
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Aug 15, 2026