Evaluate and improve AI-generated code while building verification systems for cutting-edge language models. This flexible contractor role is open to software engineers in the US, Canada, and select Western European countries.
Coding & Software
Remote
10 countries
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open to applicants in
United States Canada Austria Belgium France Germany Netherlands Switzerland Luxembourg Ireland
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps people discover opportunities, create a professional profile, and apply to projects that shape how modern AI systems are built.
Free account creation
Remote opportunities across the AI training industry
A place to build experience and grow an AI training career
About AI Code Training
Large language models learn to write better software through carefully prepared examples, expert solutions, and rigorous evaluations. In this growing field, software engineers help identify weaknesses in AI-generated code and create the feedback systems that improve model performance.
Work directly on datasets used to train and benchmark AI models
Apply software engineering judgment to evaluate model-generated solutions
Help shape how AI systems handle real software engineering tasks
The Role
OpenTrain AI is seeking an LLM Code Evaluation Software Engineer to create datasets for training, benchmarking, and advancing large language models. You will work closely with researchers to curate code examples, develop precise solutions, correct code, and assess AI-generated software across multiple programming languages.
This is a contractor, part-time engagement with flexible scheduling. The role requires a minimum commitment of 10 hours per week and may offer up to 40 hours per week. The initial duration is one month, with potential extensions based on performance and fit.
Contractor position
Flexible schedule
Minimum 10 hours per week; engagement details allow up to 40 hours per week
Initial one-month term with potential extension
No medical or paid leave included
What You'll Do
You will combine hands-on software engineering with structured evaluation of language model output. Your work will support reliable, scalable AI coding solutions and help researchers understand model capabilities throughout the software engineering cycle.
Curate code examples and build solutions in Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go.
Evaluate and refine AI-generated code for efficiency, scalability, and reliability.
Collaborate with cross-functional teams to improve AI coding solutions against industry benchmarks.
Build agents that verify code quality and identify recurring error patterns.
Form hypotheses about software engineering workflows and evaluate model capabilities at different steps.
Design verification mechanisms that automatically validate solutions to software engineering tasks.
Requirements
This role is designed for an experienced software engineer who can assess code deeply and explain technical judgments clearly. Strong full-stack development and production software experience are essential.
At least three years of software engineering experience
Deep understanding of software architecture, design, development, and debugging
Strong ability to assess code quality and conduct code reviews
Excellent oral and written communication skills
Ability to provide clear, structured evaluation rationales
Helpful Background
Experience with large language model evaluation or AI model training is helpful, but the core focus is strong software engineering judgment and the ability to evaluate coding solutions.
LLM evaluation experience is a plus
AI model training experience is a plus
Professional experience across multiple programming languages is relevant to the work
Location and Eligibility
Candidates must be based in the United States, Canada, or an eligible Western European country. The role is available in English.
United States
Canada
Austria, Belgium, France, Germany, the Netherlands, Switzerland, Luxembourg, or Ireland
English-language work
Why Join AI Training Work
AI training is the human side of building artificial intelligence. Engineers, writers, reviewers, and other specialists provide the examples and feedback that help AI systems become more capable, accurate, and useful.
Remote work with flexible hours
A chance to contribute to cutting-edge AI development
Work that can fit around other professional commitments
An opportunity to build experience in a fast-growing technology field
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.
Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.