Help train and benchmark large language models by evaluating, correcting, and verifying code across Python, JavaScript, C/C++, Java, Rust, and Go. This flexible contractor role is open to software engineers in the US, Canada, and Western Europe.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover projects, build a professional profile, and apply to opportunities that shape the future of artificial intelligence. Creating an OpenTrain account is free.
About AI Training Work
AI training is the human side of building modern artificial intelligence. Software engineers contribute by creating examples, reviewing model-generated code, rating results, and developing reliable ways to test whether AI systems can solve real engineering problems.
This work offers a direct opportunity to influence cutting-edge language models while working remotely and flexibly. Contributors with specialized technical expertise can take on projects that require advanced knowledge of software development and code quality.
- Help improve how large language models understand and generate software.
- Work remotely with a flexible schedule within the engagement requirements.
- Apply your engineering judgment to training, benchmarking, and evaluation data.
The Role
OpenTrain is hiring a contractor LLM Code Evaluation Software Engineer to help create datasets for training, benchmarking, and advancing large language models. You will work with researchers and cross-functional teams to curate code examples, develop precise solutions, correct AI-generated code, and evaluate model capabilities across software engineering tasks.
The role focuses on producing dependable coding solutions and designing verification systems that can identify errors and assess quality. Prior experience with LLM evaluation or AI model training is helpful but not required.
- Contractor, part-time engagement
- Flexible schedule with a stated minimum of 10 hours and up to 40 hours per week; structured details indicate 20+ hours per week
- Initial duration of 1 month, with potential extensions based on performance and fit
- Candidates must be based in the US, Canada, or listed Western European countries
- English-language work
What You'll Do
- Curate code examples and build solutions in Python, JavaScript including ReactJS, C/C++, Java, Rust, and Go.
- Evaluate and refine AI-generated code for efficiency, scalability, and reliability.
- Collaborate with cross-functional teams to improve AI coding solutions against industry benchmarks.
- Build agents that verify code quality and identify recurring error patterns.
- Form hypotheses about steps in the software engineering cycle and evaluate model capabilities at those steps.
- Design verification mechanisms that automatically assess solutions to software engineering tasks.
- Provide precise, clear, and structured rationales for evaluation decisions.
Requirements
This role requires several years of software engineering experience, with at least 3 years indicated in the requirements. You should be comfortable building and deploying production-grade software and explaining technical judgments clearly in both writing and conversation.
- At least 3 years of software engineering experience.
- Strong expertise building full-stack applications.
- Experience deploying scalable, production-grade software.
- Deep understanding of software architecture, design, development, and debugging.
- Strong ability to assess code quality and conduct or interpret code reviews.
- Excellent oral and written communication skills for clear evaluation rationales.
Who Should Apply
This opportunity is designed for software engineers who want to apply their development expertise to AI training and evaluation. It may be a strong fit if you enjoy analyzing how systems work, identifying error patterns, designing automated checks, and translating technical reasoning into consistent evaluation criteria.
- Software engineers with full-stack and production deployment experience.
- Developers who work with one or more of Python, JavaScript, ReactJS, C/C++, Java, Rust, or Go.
- Engineers interested in benchmarking and improving large language models.
- Candidates with previous LLM evaluation or AI model training experience, which is a plus.
- Applicants based in the United States, Canada, Austria, Belgium, France, Germany, or other eligible Western European countries.
Engagement Details
This is a flexible, part-time contractor engagement. The initial term is one month, with possible extensions based on performance and fit. The role does not include medical or paid leave.
- Employment type: Contractor and part-time
- Location: Remote, limited to eligible countries
- Duration: 1 month initially, with potential extension
- Schedule: Flexible, up to 40 hours per week
- Compensation: USD terms are listed, but no rate is specified in the provided details
- Benefits: No medical or paid leave
How to Apply Through OpenTrain
Create a free OpenTrain account, build your profile around your software engineering experience, and apply in minutes. OpenTrain helps contributors start and grow careers in the fast-moving AI training and data-labeling industry, where human technical judgment remains essential to building capable and reliable AI systems.
- Highlight your full-stack development and production deployment experience.
- List the programming languages and frameworks you use confidently.
- Describe your experience with debugging, code review, architecture, or automated verification.
- Mention LLM evaluation or AI training experience if applicable.