Code Generation Model Evaluation Engineer
Use advanced programming skills to create coding tasks, evaluate AI-generated solutions, and improve real-world software across Python, Java, Rust, C++, Go, and TypeScript.
Posted Sep 14, 2026
Join a remote, part-time project evaluating AI code generation and improving real-world software tasks. Use Python, Java, Rust, C++, Go, or TypeScript expertise for $50-$100/hr in an under-20-hour weekly commitment.
Coding & Software
$50–$100/hr
Compensation
8 countries
Eligibility
Intermediate
Experience
Sep 16, 2026
Posted
Open to applicants in
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is the hiring and contracting organization for this role, connecting experienced technical contributors with projects that help shape how modern AI systems learn and perform.
Creating an OpenTrain account is free, and selected contractors can apply their existing software engineering expertise to cutting-edge AI work without needing previous AI industry experience.
AI training is the human side of building artificial intelligence. Technical contributors create, review, and evaluate examples that help AI models generate better code, reason through complex problems, and produce more accurate outputs.
In this project, your software engineering knowledge will support code generation workflows and model evaluation. Your analysis, coding tasks, annotations, and feedback will help assess how well AI systems handle realistic programming challenges.
OpenTrain AI is seeking Research Engineers for a project focused on code generation and model evaluation. You will analyze and improve software across multiple programming languages, create and review coding tasks, and provide precise technical feedback that supports AI model training and assessment.
This is an intermediate-level contractor role for candidates with software engineering and AI code generation subject matter expertise. The work requires strong competitive programming and coding problem analysis skills, careful code validation, and the ability to explain technical decisions clearly in English.
You will work across diverse codebases and contribute to the development of high-quality programming examples for AI evaluation. Assignments may involve debugging existing implementations, building features, reviewing alternative solutions, and documenting decisions for other contributors and technical stakeholders.
Candidates should bring substantial software engineering ability and experience analyzing coding problems. You must be comfortable working independently in a remote, collaborative environment and producing well-documented, consistent work under tight deadlines.
Expertise in at least one required programming language is expected, along with solid knowledge of algorithms and data structures. Open-source contributions or participation in collaborative software projects is also required.
The listed pay rate is $50-$100 per hour. Compensation is output-based, with experts paid per task that meets project specifications. The time required can vary based on experience and workflow, and minimum submission requirements apply, including a minimum number of tasks each week.
Roles are typically filled within 48 hours. If selected, you should be ready to begin your first tasks within 24-48 hours after completing onboarding.
Every major AI system depends on people who can prepare, review, and improve the examples used for training. By bringing your engineering judgment to code generation and model evaluation, you can directly influence how AI systems solve programming problems and communicate their results.
OpenTrain makes it easier to start and grow a career in AI training and data labeling. Create a free profile and apply to this technical contractor opportunity in minutes.
Keep exploring
Use advanced programming skills to create coding tasks, evaluate AI-generated solutions, and improve real-world software across Python, Java, Rust, C++, Go, and TypeScript.
Posted Sep 14, 2026
Create and validate challenging engineering simulation benchmarks that train and evaluate AI agents. This remote contractor role combines advanced engineering design, Python, open-source simulation, and model failure analysis.
Posted Sep 11, 2026
Review and benchmark AI-generated software as a remote contractor in the United States. Use your software engineering expertise to test code, assess explanations, investigate failures, and improve coding evaluation standards.
Posted Jul 17, 2026
Browse related job pages
Expertise
Locations
Languages