Skip to content
OpenTrain AIFor AI Companies

AI Agent Task Mining Software Engineer

Apply now

AI Agent Task Mining Software Engineer

Build realistic, multi-step tasks that train and evaluate AI agents performing enterprise work. This part-time contract is suited to Python backend engineers who can turn complex workflows into precise objectives and rubrics.

OpenTrain AI

Coding & Software

Remote

9 countries

Eligibility

Entry

Experience

Sep 12, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover opportunities, build a professional profile, and grow durable experience in a fast-moving field.

As an OpenTrain contractor, you will contribute directly to the development and evaluation of advanced AI systems while building a portfolio of specialized work.

  • Contractor and part-time opportunity
  • Open to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, and Mexico
  • English-language work
  • Time commitment of 20 or more hours per week

About AI Agent Training

AI training is the human side of building artificial intelligence. People create examples, evaluate model behavior, and define what high-quality outcomes look like so AI systems can become more accurate, useful, and reliable.

In this role, you will work on the software engineering and evaluation side of that process by modeling authentic enterprise workflows and turning them into challenging, measurable tasks for AI agents.

  • Work at the intersection of backend engineering and AI evaluation
  • Help shape how AI agents perform realistic digital knowledge work
  • Contribute to a rapidly growing area of technology through flexible contract work

The Role

OpenTrain is seeking an AI Agent Task Mining Software Engineer to research real-world enterprise workflows and convert them into rigorous training and evaluation tasks. You will create realistic, long-horizon objectives, define how successful completion should be judged, and ensure that each task is achievable and accurately represented in simulated business-application environments.

The role requires strong Python backend engineering, sound engineering judgment, and the ability to explain expected outcomes with precise evaluation rubrics. You will also investigate failures and improve the quality standards used to assess agent performance.

  • Primary focus: AI agent task mining and evaluation
  • Core role nouns: software engineer, task author, evaluator, and QA contributor
  • Experience level listed for the opportunity: entry level
  • Part-time contractor engagement requiring 20 or more hours per week

What You'll Do

You will combine workflow research, backend engineering, task design, and quality assurance. The work involves validating tasks from end to end and ensuring that objectives, environments, and grading criteria accurately reflect realistic enterprise work.

  • Research real data, tools, and workflows to understand knowledge work across enterprise applications.
  • Author long-horizon agent tasks that mirror genuine digital work and support AI training and evaluation.
  • Write precise rubrics describing correct, complete, and high-quality outcomes.
  • Review tasks for realism, correctness, scope, ambiguity, missing edge cases, and grading gaps.
  • Validate objectives end to end against connector environments that replicate Slack, Linear, Jira, Notion, Gmail, and wikis.
  • Reproduce and debug failures, communicate specific findings, and improve task quality.
  • Improve quality-control standards, checklists, and related processes.

Required Skills

This role is designed for a strong generalist backend engineer who can reason carefully about real-world software workflows and communicate technical judgments clearly. You should be comfortable identifying subtle problems in task definitions, simulated environments, and evaluation criteria.

  • Strong backend software engineering experience, primarily with Python
  • Ability to model real-world enterprise workflows as precise, multi-step agent tasks
  • Skill in writing rigorous evaluation rubrics and clear QA findings
  • Sound judgment for identifying ambiguity, edge cases, and unrealistic assumptions
  • Working knowledge of GCP, Docker, virtual machines, and Harbor
  • Excellent written communication, especially for rubric and QA writing
  • Daily proficiency with AI coding tools such as Claude Code, Cursor, or Copilot

Helpful Background

The following experience is helpful but is not listed as required. Strong backend engineers who can also contribute to connector-side work may be especially well suited to the opportunity.

  • Building or evaluating agentic systems or large language model systems
  • Working with SaaS APIs and data models
  • Contributing to data annotation, evaluation design, or quality assurance at scale
  • Contributing to connector-side engineering work

Why This Work Matters

Every major AI system depends on people who prepare examples, test behavior, and define quality. By translating authentic enterprise work into structured agent tasks, you will help make AI systems more capable of completing useful multi-step workflows accurately and reliably.

AI training and data-labeling work can provide flexible, remote opportunities for people with technical or domain expertise. OpenTrain helps contributors build a profile and turn this experience into a longer-term AI training portfolio.

  • Work on cutting-edge AI agent capabilities
  • Apply practical backend engineering to emerging AI systems
  • Build credible experience in task design, evaluation, and quality assurance

How to Apply

Review the opportunity through OpenTrain and apply with details that demonstrate your Python backend engineering, enterprise workflow modeling, rubric writing, and AI coding tool experience. Be prepared to show how you identify ambiguity, edge cases, unrealistic assumptions, and gaps in evaluation quality.

  • Apply through OpenTrain AI
  • Highlight relevant Python backend engineering experience
  • Describe experience with GCP, Docker, virtual machines, and Harbor
  • Mention daily use of Claude Code, Cursor, Copilot, or comparable AI coding tools
  • Confirm your availability for 20 or more hours per week

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Agent Task QC Engineer

Help improve frontier AI agents by reviewing long-horizon tasks, testing rubrics, debugging failures, and catching edge cases. This flexible contractor role requires Python, cloud tooling, English fluency, and at least 20 hours per week.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +5 more
English
Part-time · Flexible
Entry level

Posted Sep 12, 2026

Software Engineering AI Task Author

Author realistic debugging, integration, and configuration tasks that train and evaluate advanced AI agents. This part-time contractor role offers $70-$120 per hour for eligible contributors in India.

Coding & Software
Computer Code Programming
Remote · India
English
Part-time · Flexible
Entry level
Hourly · $70–$120/hr

Posted Aug 18, 2026

Senior Coding-Agent Benchmark Engineer

Create and evaluate realistic software-engineering benchmarks for coding agents using production-like repositories, secure coding, and rigorous testing. This fully remote contractor assignment runs 4 to 8 weeks.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 2, 2026