Evaluate AI agents on realistic software engineering tasks and build reliable MCP-based reinforcement learning environments. This flexible remote contractor role offers an ideal rate range of $60 to $120 per hour.
Coding & Software
100% Remote Hourly · $60–$120/hr
$60–$120/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Aug 25, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build a credible profile, and grow long-term experience teaching and evaluating AI systems. Creating an OpenTrain account is free.
Remote contractor opportunity
Flexible schedule across available days
Approximately 15 hours per week
Work available worldwide
About AI Software Evaluation
AI training is the human side of building artificial intelligence. Engineers and other specialists create tasks, review model behavior, and provide evidence-based feedback that helps advanced systems become more capable, reliable, and useful. In this role, your software engineering expertise will directly support the evaluation of AI agents.
Work on cutting-edge AI training and evaluation
Assess how agents reason through realistic development tasks
Help improve model behavior through rigorous technical feedback
The Role
OpenTrain is recruiting an MCP AI Software Evaluation Engineer to evaluate AI systems on complex software engineering tasks. You will create reinforcement learning environments that test how agents use Model Context Protocol tools to discover information, reason through problems, and complete realistic development work.
Assignments may involve bug fixing, feature implementation, codebase refactoring, performance optimization, and designing reliable evaluation methods. Prior AI training experience is not required; strong real-world software engineering expertise is the most important qualification.
Evaluate AI agents interacting with MCP tools and MCP servers
Support assessments of complex programming and software engineering work
Work remotely as a flexible, part-time contractor
What You'll Do
You will build reproducible environments, deterministic verification, and golden reference solutions for software engineering evaluations. You will compare agent implementations and problem-solving approaches against clear project guidelines, then document evidence-based judgments so evaluation results can improve model behavior.
Design reproducible environments for software engineering evaluations
Create deterministic verification methods and golden reference solutions
Assess bug fixes, feature development, refactoring, and performance improvements
Evaluate how effectively AI agents use MCP tools
Review implementation quality and debugging approaches
Maintain consistency and quality across repeated tasks
Communicate technical feedback clearly in a remote contractor setting
Required Qualifications
You should be proficient in at least one of C++, Python, Java, Go, TypeScript, or Rust, with strong knowledge of algorithms and data structures. Demonstrated experience delivering maintainable software, developing features, debugging complex issues, refactoring codebases, and optimizing systems for performance and scalability is required.
Familiarity with Model Context Protocol tools and agent interaction with MCP servers is valuable. Experience with large-scale or distributed codebases, rigorous code reviews, software best practices, or modern AI and machine learning systems is helpful. Strong written and verbal communication, careful attention to detail, and the ability to follow detailed technical guidelines are essential.
Proficiency in C++, Python, Java, Go, TypeScript, or Rust
Strong understanding of algorithms and data structures
Experience debugging complex software issues
Experience developing features and refactoring codebases
Experience optimizing performance and scalability
Ability to design reproducible environments and deterministic verification
Ability to create golden reference solutions for software evaluations
Strong written communication and attention to detail
Ability to collaborate effectively and follow detailed technical guidelines
Schedule And Compensation
This is a remote contractor opportunity requiring approximately 15 hours per week, with flexible scheduling across available days. The listed ideal rate range is $60 to $120 per hour. Confirm the applicable compensation terms for assigned work.
Contractor and part-time engagement
Approximately 15 hours per week
Flexible scheduling
Ideal rate range: $60 to $120 per hour
Worldwide opportunity
English-language work
Build Your AI Training Career
OpenTrain gives freelancers one place to manage AI training opportunities, demonstrate relevant experience, and develop a durable AI training portfolio. Your software engineering background can help shape how state-of-the-art AI systems solve programming and development problems.
Build a profile showcasing your AI training experience
Discover projects aligned with your technical skills
Apply through OpenTrain and grow your long-term portfolio
Help build reinforcement learning environments that evaluate AI agents on complex software engineering tasks using MCP tools. This flexible remote contract offers approximately 15 hours per week at $60–$120/hour.
Use your software engineering expertise to build MCP-powered reinforcement learning environments that test how AI agents solve real-world coding problems. This flexible contractor role pays $60 to $120 per hour.
Build reproducible reinforcement learning environments that evaluate how AI models solve complex software engineering problems with MCP tools. This flexible, worldwide contract role pays $80 to $120 per hour for less than 20 hours weekly.