Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free, and this opportunity gives experienced software engineers a way to contribute directly to the development of advanced AI systems.
Contractor position
Part-time engagement
Worldwide opportunity
English-language work
20+ hours per week
About AI Training and LLM Evaluation
AI training is the human side of building artificial intelligence. People prepare examples, write and review responses, and evaluate model outputs so that modern AI systems become more accurate, reliable, and useful.
In this role, your software engineering expertise will help train and benchmark large language models that generate code. You will assess their capabilities across software engineering tasks and help create verification methods that measure quality, efficiency, and reliability.
Work on cutting-edge large language model development
Shape how AI systems generate and evaluate software
Contribute remotely with flexible part-time hours
The Role
OpenTrain AI is recruiting a Senior Software Engineer for LLM Evaluation. In this contractor role, you will create and refine high-quality datasets for training and benchmarking large language models.
You will curate code examples, build and correct solutions in multiple programming languages, and evaluate AI-generated code for efficiency and reliability. You will also work with cross-functional teams to improve AI-driven coding solutions against industry performance benchmarks.
Subject matter: LLM code evaluation and dataset curation
Primary focus: software engineering evaluation and programming data
Listed experience level: entry level, with substantial professional experience required in the qualifications
What You’ll Do
Your work will combine hands-on software development with structured evaluation of AI-generated solutions. You will need to explain your judgments clearly and design processes that can assess software engineering performance consistently.
Curate code examples for AI model training initiatives.
Build and correct solutions in Python, JavaScript including ReactJS, C/C++, Java, Rust, and Go.
Evaluate and refine AI-generated code for efficiency, scalability, and reliability.
Collaborate with cross-functional teams to improve AI-driven coding solutions against industry performance benchmarks.
Build agents that verify code quality and identify error patterns.
Hypothesize about steps in the software engineering cycle and evaluate model capabilities at each step.
Design verification mechanisms that can automatically verify solutions to software engineering tasks.
This role requires several years of software engineering experience, including at least two years of continuous full-time experience at a top-tier product company. The listed examples include Google, Stripe, Amazon, Apple, Meta, Netflix, and Microsoft.
The role also calls for strong full-stack development experience and the ability to deploy scalable, production-grade software using modern languages and tools.
Several years of software engineering experience.
At least 2 years of continuous full-time experience at a top-tier product company such as Google, Stripe, Amazon, Apple, Meta, Netflix, or Microsoft.
Strong expertise building full-stack applications.
Deep understanding of software architecture, design, development, debugging, and code quality review.
Excellent oral and written communication skills.
Ability to write clear, structured evaluation rationales.
Experience with Python, JavaScript, ReactJS, C/C++, Java, Rust, and/or Go is essential for success.
Who Should Apply
This opportunity is suited to software engineers who can combine deep technical judgment with careful written evaluation. It may be especially relevant to engineers who have built full-stack systems, reviewed production code, debugged complex software, or worked with several of the listed programming languages.
You should be comfortable assessing not only whether code works, but also whether it is scalable, efficient, reliable, and aligned with sound software engineering practices.
Experienced full-stack software engineers
Engineers with strong architecture and code review backgrounds
Developers comfortable evaluating code across multiple languages
Clear technical communicators who can justify structured judgments
Candidates able to commit to at least 20 hours per week
How the Engagement Works
This is a worldwide, part-time contractor opportunity conducted in English. The time requirement is 20 or more hours per week, making it suitable for professionals seeking flexible remote work while contributing to the development of advanced AI coding systems.
Apply through OpenTrain to create your profile and submit your application. OpenTrain accounts are free, and the platform supports contributors as they build careers in AI training and data labeling.
Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.
Evaluate how AI models fix real software bugs in open-source Python repositories. This flexible, part-time contractor role focuses on GitHub issue triage, Docker environments, testing, and LLM performance assessment.
Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.