Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Aug 30, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain connects experienced technical contributors with projects that help build and evaluate modern artificial intelligence systems.
Creating an OpenTrain account is free, and candidates can build a profile and apply in minutes. This contract offers the opportunity to combine practical software engineering with work at the cutting edge of AI development.
About AI Training and LLM Evaluation
AI training is the human side of building artificial intelligence. People prepare datasets, review model outputs, and test whether AI systems can complete realistic tasks accurately and reliably.
In this role, you will help evaluate large language models, or LLMs, on software engineering problems. By testing models against real codebases and bug-fixing scenarios, you will contribute to datasets that help AI systems work more effectively with code.
The Role
OpenTrain is seeking an experienced software engineer with strong C++ skills to build LLM evaluation and training datasets for realistic software engineering problems. You will work with public GitHub repositories to create verifiable software engineering tasks, assess how well LLMs understand and fix code, and expand dataset coverage across programming languages and difficulty levels.
The role is fully remote and is available to contractors in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico. The assignment is expected to last one month, with an anticipated start next week.
Role type: Contractor and part-time
Experience level listed: Entry level
Language: English
Structured time requirement: 20+ hours per week
Role schedule: At least 40 hours per week with four hours of overlap with PST
Duration: One month
Expected start: Next week
What You'll Do
You will investigate real-world software repositories and create reliable evaluation tasks that show whether an LLM can reason about, modify, and test code. The work combines repository analysis, local development, test evaluation, and collaboration with AI researchers.
Analyze and triage GitHub issues across trending open-source libraries.
Set up and configure code repositories, including Dockerization and environment setup.
Evaluate unit test coverage and quality.
Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.
Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.
Take the opportunity to lead a team of junior engineers.
Requirements
This opportunity requires hands-on software engineering experience and the ability to work confidently inside complex, real-world codebases. You should be comfortable setting up projects locally, changing code, and running tests to verify behavior.
At least three years of overall software engineering experience.
Strong experience with C++.
Proficiency with Git, Docker, and basic software pipeline setup.
Ability to understand and navigate complex codebases.
Comfort running, modifying, and testing real-world projects locally.
Familiarity with LLM evaluation or research is helpful but not required.
Who Should Apply
This role may suit a C++ engineer who enjoys debugging, developer tooling, open-source software, and testing the limits of AI systems. It is especially relevant for engineers who have worked on developer tools, automation agents, or previous LLM research and evaluation projects.
The work is fully remote, making it possible to contribute from an eligible country while working on practical problems that influence how language models interact with real code.
Software engineers with strong C++ experience.
Engineers who can independently configure and run unfamiliar repositories.
Contributors interested in AI research, LLM evaluation, or coding datasets.
Candidates comfortable collaborating with researchers and potentially leading junior engineers.
How to Apply
Apply through OpenTrain to be considered for this remote software engineering evaluation contract. Review the location, schedule, and technical requirements carefully before applying, especially the requested PST overlap and the differing time markers in the listing.
OpenTrain helps contributors discover and grow in AI training work, from code evaluation and model testing to other projects that shape how advanced AI systems are built.
Eligible locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.
Work arrangement: Fully remote.
Apply through OpenTrain and create a free contributor profile.
Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.
Help improve AI-assisted software development by evaluating LLMs on real Go codebases, open-source issues, and bug-fixing tasks. Lead related projects while working remotely for 20+ hours per week.