Skip to content
OpenTrain AIFor AI Companies

Software Engineering AI Evaluation Engineer

Build reproducible coding evaluation environments and golden solutions that help improve AI models on complex software engineering tasks. This remote contractor role offers flexible project work with compensation metadata indicating $50–$150 per hour.

OpenTrain AI

Coding & Software

100% Remote Hourly · $50–$150/hr

$50–$150/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Jul 12, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover projects, build a professional profile, and apply for opportunities in a fast-growing field where people help shape how modern AI systems work.

  • Create a free OpenTrain account and build your AI training profile
  • Discover remote projects aligned with your technical experience
  • Develop a portfolio of work supporting the next generation of AI

About AI Training and Software Evaluation

AI training is the human side of building artificial intelligence. For software-focused models, expert contributors create realistic coding tasks, assess model-generated solutions, and document the reasoning needed to distinguish robust engineering from unreliable output.

  • Apply software engineering expertise to cutting-edge AI development
  • Work remotely with a flexible contractor arrangement
  • Contribute technical examples and evaluations that improve model performance

The Role

OpenTrain is seeking a Software Engineering AI Evaluation Engineer to create reinforcement learning environments that evaluate how AI models solve complex software engineering problems. You will convert real engineering challenges into reproducible tasks with golden reference solutions and provide expert review across a range of codebases.

The work includes code fixes, feature development, legacy-system refactoring, and performance improvement. Previous AI experience is helpful but not required.

  • Remote contractor engagement
  • Part-time project work
  • Approximately 15 hours per week according to the role description
  • Structured engagement metadata indicates a time requirement of 20+ hours per week
  • English-language work
  • Worldwide opportunity

What You'll Do

You will combine hands-on software engineering, debugging, technical reasoning, and clear documentation to produce high-quality AI training data. Your work will help establish reliable ways to measure model performance on realistic engineering problems.

  • Create reproducible environments for testing AI performance on software engineering problems
  • Develop golden reference solutions for code fixing, feature implementation, refactoring, and optimization tasks
  • Contribute expert code samples, debugging strategies, and development insights
  • Analyze defects and performance bottlenecks
  • Implement robust and maintainable solutions
  • Refactor legacy code to improve clarity, efficiency, scalability, and reliability
  • Document technical reasoning, solution approaches, and code decisions
  • Review peer-contributed code and technical submissions for accuracy and clarity

Required Skills and Experience

This role requires significant hands-on expertise in at least one of the listed programming languages, together with strong fundamentals in algorithms, data structures, and software engineering. You should be comfortable working through complex systems and explaining your technical decisions clearly.

  • Proficiency in at least one of Python3, Java, Rust, Go, C++, or TypeScript
  • Strong understanding of algorithms, data structures, and core software engineering principles
  • Ability to debug complex software defects and performance bottlenecks
  • Experience delivering effective optimizations
  • Ability to refactor and modernize large or legacy codebases
  • Experience delivering high-impact features from conception through completion
  • Clear technical communication and documentation skills
  • Ability to design reproducible reinforcement learning environments and golden reference solutions

Compensation and Engagement

This is a part-time contractor engagement. Source compensation metadata indicates an ideal range of $50–$150 per hour, with a listed maximum hourly rate of $150. Project terms also describe output-based task compensation, so candidates should review the applicable terms before beginning work.

  • Payment type: hourly, with project terms also describing output-based compensation
  • Compensation range indicated in the source metadata: $50–$150 per hour
  • Employment type: contractor and part time
  • Work arrangement: remote and worldwide

Why Build Your AI Training Career with OpenTrain

OpenTrain gives you one place to discover AI training opportunities, showcase credible experience, and build a portfolio you control. Whether you are expanding from software engineering into AI evaluation or developing a long-term technical freelance career, your OpenTrain profile can help you present your skills and find relevant projects.

  • Work on the human side of modern AI development
  • Build evidence of specialized software engineering evaluation experience
  • Find flexible remote opportunities in a rapidly growing industry
  • Create and strengthen your profile at no cost

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Software Engineering Code Evaluator

Evaluate and improve AI-generated code while creating benchmarking datasets for advanced software engineering models. This expert, remote contract role is part time at less than 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 16, 2026

Software Engineer for AI Training Environments

Join OpenTrain AI as a part-time contractor building reinforcement-learning environments and reproducible software tasks that evaluate model programming ability — remote, worldwide, under 20 hrs/week, paying up to $150/hr.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$150/hr

Posted Jul 23, 2026

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026