OpenTrain is the leading platform for people building careers in AI training and data labeling. We connect skilled contributors with meaningful evaluation and annotation work, enabling you to build a portfolio, find flexible remote projects, and grow into a durable freelance career in a fast-growing industry.
For this role, OpenTrain AI is the hiring and contracting organization. You will work directly with our project teams to evaluate model outputs, improve benchmarks, and help shape how AI systems are measured and improved.
About AI Training Work
AI training (also called data labeling or human feedback work) is the human side of building modern AI: people prepare, review, and grade examples that models learn from. Contributors do tasks like annotating text, images, or code; rating and ranking model outputs; and refining benchmarks that guide model development.
This work is 100% remote, often flexible and part-time, and accessible to many contributors—while specialist tasks such as code evaluation pay more for software engineering experience. By joining OpenTrain you help shape cutting-edge models while keeping flexible hours.
The Role
We are recruiting an AI Code Evaluation Engineer to assess and benchmark the coding capabilities of frontier AI models. You will evaluate AI-generated code for correctness, quality, and adherence to requirements, reproduce and debug issues, and help create high-quality evaluation datasets and rubrics.
This position suits engineers who enjoy code review, debugging, problem-solving, and applying strong software engineering judgment to technical tasks.
What You'll Do
Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements.
Debug code, reproduce issues, and verify fixes across different programming environments.
Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy.
Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
Identify edge cases, failure modes, and areas where AI systems struggle with software engineering problems.
Document findings clearly and provide structured feedback to improve evaluation quality and consistency.
Collaborate with project teams to establish quality standards and evaluation methodologies.
Requirements
You must meet all of the stated requirements below to be considered for this role.
Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical field.
3+ years of professional software engineering experience.
Strong proficiency in one or more of: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
Strong understanding of data structures, algorithms, software design principles, and debugging methodologies.
Experience performing code reviews and evaluating code quality in production or large-scale codebases.
Familiarity with version control systems such as Git and modern software development workflows.
Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
Strong written communication skills and attention to detail.
Helpful Background
The following experiences are not required but will help you excel in this role:
Experience with AI/ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects.
Experience evaluating AI-generated code, benchmark creation, or software quality assessment.
Schedule, Location, and Contract Details
This is a remote contractor role available to applicants located in the United States only.
Time commitment: at least 4 hours per day with a minimum of 20 hours per week. You must have 4 hours of overlap with Pacific Standard Time (PST). Contract duration: 1 month. Employment types: contractor, part-time.
Location: US only (remote).
Minimum weekly commitment: 20 hours (at least 4 hours/day).
PST overlap: 4 hours required.
Contract length: 1 month.
Who Should Apply and How It Works
Apply if you are an experienced software engineer who enjoys reviewing code, diagnosing tricky bugs, and formalizing evaluation criteria. You will help determine how well AI systems perform on real engineering tasks and contribute to benchmarks that guide model improvement.
OpenTrain makes it easy to build an AI training portfolio and find flexible remote projects. If this role fits your skills and availability, create or update your OpenTrain profile and apply — our team will review your qualifications and follow up with next steps.
Design and validate multi-agent coding benchmarks using real open-source code changes, Docker, and Python verification scripts. Remote 4-week contractor role for developers in select countries, requiring daily overlap with PST.
Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.
Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.