Skip to content
OpenTrain AIFor AI Companies

C++ LLM Evaluation Software Engineer

Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.

OpenTrain AI

Coding & Software

Remote

9 countries

Eligibility

Entry

Experience

Jul 17, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people building careers in AI training and data labeling. We help contributors discover projects, build a unified AI-training portfolio, and grow into durable freelance careers working on the human side of AI.

  • We hire and contract engineers and annotators to create and evaluate training data that shapes state-of-the-art AI.
  • This role is directly with OpenTrain AI as the hiring organization.

About AI training and LLM evaluation

AI training (data labeling/annotation) is the human work behind modern models: designing tasks, evaluating outputs, and creating examples that teach models to understand code, language, images, and audio. Evaluating LLMs against real engineering problems helps improve model correctness and reliability.

  • This role sits at the intersection of software engineering and LLM evaluation—you'll use real repositories and tests to measure model performance.
  • Contributors often work remotely, part time, and can gain hands-on experience in cutting-edge LLM research and tooling.

Role overview

You will create verifiable software engineering evaluation tasks from public repository histories, triage issues, configure environments, run and modify code, and rate LLM outputs against real bug-fix and test scenarios. Expect close collaboration with researchers to select challenging repositories and design meaningful evaluations.

  • Employment: Contractor, part-time.
  • Time commitment: Flexible options — 20, 30, or 40 hours per week.
  • Work model: Fully remote; English required.

What you'll do

Practical, hands-on engineering and evaluation work focused on C++ codebases and LLM behavior. The role combines triage, environment setup, test execution, and qualitative evaluation.

  • Analyze and triage GitHub issues across trending open-source C++ libraries.
  • Set up and configure repositories, including Dockerization and environment setup.
  • Evaluate unit test coverage and test quality to design evaluation criteria.
  • Modify and run code locally to create reproducible bug-fix scenarios for LLMs.
  • Collaborate with researchers to identify hard-to-solve repositories and craft evaluation tasks.
  • Opportunity to lead and mentor a small team of junior engineers on evaluation projects.

Requirements

Candidates must meet the core technical requirements below; this role expects practical experience running and debugging real projects and familiarity with development tooling.

  • Minimum 3+ years of software engineering experience.
  • Strong experience with C++ and working within complex codebases.
  • Proficiency with Git, Docker, and basic software pipeline setup.
  • Comfortable running, modifying, and testing real-world projects locally.
  • Experience contributing to or evaluating open-source projects is a plus.
  • Language: English (required).
  • Eligible countries: IN, PK, NG, KE, EG, GH, BD, TR, MX (role is not worldwide).

Helpful background

These are not strict requirements but will make you more effective from day one.

  • Previous participation in LLM research, evaluation, or RLHF projects.
  • Experience building or testing developer tools, automation agents, or CI workflows.

How we work and next steps

OpenTrain runs remote, contractor-first projects designed for flexible schedules. Compensation details are not specified in this listing and will be shared during the application process. If you meet the requirements, create an OpenTrain account and apply to be considered.

  • Contractor role with part-time hours (choose from 20/30/40 hrs/week).
  • Work involves evaluating LLM outputs (rating/evaluation) and hands-on C++ coding tasks.
  • If selected, you'll collaborate directly with researchers and engineering leads to design reproducible evaluations.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Senior Software Engineer - C# (LLM Evaluation & Repository Validation)

Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

LLM Evaluation Software Engineer (Ruby)

Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Senior LLM Code Evaluation Engineer

Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026