Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 17, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects that help shape how modern AI systems learn, reason, and perform.
This role offers the opportunity to build a durable freelance career while gaining hands-on experience in LLM evaluation and AI training data creation.
Fully remote contractor opportunity
Flexible commitment of 20, 30, or 40 hours per week
Available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico
About AI Training and LLM Evaluation
AI training is the human side of building artificial intelligence. Engineers and other specialists create, review, and evaluate examples that help models produce more reliable results. In this project, your software engineering judgment will be used to test how well LLMs understand, modify, and debug real-world code.
Work on cutting-edge evaluation of generative AI
Create realistic examples based on open-source software
Help researchers identify the code problems that challenge LLMs most
The Role
OpenTrain is recruiting a software engineering contractor to build and evaluate LLM datasets focused on software development. You will create verifiable engineering tasks from public repository histories, triage GitHub issues, configure development environments, and evaluate LLM performance on real code problems.
The role requires strong C++ experience and the ability to navigate complex codebases. You will work with researchers to design tasks that push the boundaries of LLM code understanding, with an opportunity to lead a team of junior engineers.
Core focus: C++ LLM code evaluation
Contractor and part-time engagement
Minimum commitment: 20 hours per week
What You'll Do
You will work directly with real repositories and software engineering scenarios rather than simplified coding exercises. Your evaluations will help determine whether LLMs can understand issues, make appropriate changes, and produce working solutions.
Analyze and triage GitHub issues across trending open-source libraries
Set up and configure code repositories, including Dockerization and development environment setup
Evaluate unit test coverage and quality
Modify and run codebases locally to assess LLM performance in bug-fixing scenarios
Collaborate with researchers to identify repositories and issues that are challenging for LLMs
Take the opportunity to lead a team of junior engineers
Required Skills and Experience
This opportunity is listed at an entry-level project level, while the role itself requires at least three years of overall software engineering experience. Candidates should be comfortable working independently with real-world C++ projects and local development environments.
At least 3 years of overall software engineering experience
Strong experience with C++
Proficiency with Git, Docker, and basic software pipeline setup
Ability to understand and navigate complex codebases
Comfort running, modifying, and testing real-world projects locally
Experience contributing to or evaluating open-source projects is a plus
Helpful Background
Experience with AI research or developer tooling can help you contribute effectively, but the central requirement is strong practical software engineering ability.
Previous participation in LLM research or evaluation projects
Experience building or testing developer tools
Experience working with automation agents
Why Join This Project
AI training and data-labeling work is a fast-growing way to work in technology. Contributors apply their professional expertise to the datasets and evaluations behind modern AI systems, often through flexible remote arrangements that fit alongside other commitments.
Work fully remotely
Choose a 20, 30, or 40 hour weekly commitment
Build experience at the forefront of LLM evaluation
Contribute to the creation of high-quality AI training data
How to Apply
Create a free OpenTrain account to build your profile and apply in minutes. Highlight your C++ experience, work with Git and Docker, ability to navigate complex codebases, and any open-source or LLM evaluation background.
Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.
Evaluate how language models solve real software engineering problems using C#, GitHub repositories, Docker, and unit-test analysis. This remote contractor role requires 20 or more hours weekly and four hours of daily PST overlap.
Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.