Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 17, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We help contributors discover projects, build a unified AI-training portfolio, and grow into durable freelance careers working on the human side of AI.
We hire and contract engineers and annotators to create and evaluate training data that shapes state-of-the-art AI.
This role is directly with OpenTrain AI as the hiring organization.
About AI training and LLM evaluation
AI training (data labeling/annotation) is the human work behind modern models: designing tasks, evaluating outputs, and creating examples that teach models to understand code, language, images, and audio. Evaluating LLMs against real engineering problems helps improve model correctness and reliability.
This role sits at the intersection of software engineering and LLM evaluation—you'll use real repositories and tests to measure model performance.
Contributors often work remotely, part time, and can gain hands-on experience in cutting-edge LLM research and tooling.
Role overview
You will create verifiable software engineering evaluation tasks from public repository histories, triage issues, configure environments, run and modify code, and rate LLM outputs against real bug-fix and test scenarios. Expect close collaboration with researchers to select challenging repositories and design meaningful evaluations.
Employment: Contractor, part-time.
Time commitment: Flexible options — 20, 30, or 40 hours per week.
Work model: Fully remote; English required.
What you'll do
Practical, hands-on engineering and evaluation work focused on C++ codebases and LLM behavior. The role combines triage, environment setup, test execution, and qualitative evaluation.
Analyze and triage GitHub issues across trending open-source C++ libraries.
Set up and configure repositories, including Dockerization and environment setup.
Evaluate unit test coverage and test quality to design evaluation criteria.
Modify and run code locally to create reproducible bug-fix scenarios for LLMs.
Collaborate with researchers to identify hard-to-solve repositories and craft evaluation tasks.
Opportunity to lead and mentor a small team of junior engineers on evaluation projects.
Requirements
Candidates must meet the core technical requirements below; this role expects practical experience running and debugging real projects and familiarity with development tooling.
Minimum 3+ years of software engineering experience.
Strong experience with C++ and working within complex codebases.
Proficiency with Git, Docker, and basic software pipeline setup.
Comfortable running, modifying, and testing real-world projects locally.
Experience contributing to or evaluating open-source projects is a plus.
Language: English (required).
Eligible countries: IN, PK, NG, KE, EG, GH, BD, TR, MX (role is not worldwide).
Helpful background
These are not strict requirements but will make you more effective from day one.
Previous participation in LLM research, evaluation, or RLHF projects.
Experience building or testing developer tools, automation agents, or CI workflows.
How we work and next steps
OpenTrain runs remote, contractor-first projects designed for flexible schedules. Compensation details are not specified in this listing and will be shared during the application process. If you meet the requirements, create an OpenTrain account and apply to be considered.
Contractor role with part-time hours (choose from 20/30/40 hrs/week).
Work involves evaluating LLM outputs (rating/evaluation) and hands-on C++ coding tasks.
If selected, you'll collaborate directly with researchers and engineering leads to design reproducible evaluations.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.