Skip to content
OpenTrain AIFor AI Companies

LLM Evaluation & Repository Validation Engineer

Evaluate real software repositories and help measure how well AI coding systems fix bugs. This remote, three-month contractor assignment offers 20 hours per week for engineers skilled in codebases, testing, Git, and Docker.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover projects, build a professional profile, and apply for opportunities that match their skills.

OpenTrain AI is hiring and contracting for this remote software engineering evaluation assignment.

About AI Training and Software Evaluation

AI training is the human side of building artificial intelligence. For coding systems, experienced software professionals help evaluate generated code, test whether bugs are fixed correctly, and create reliable benchmarks from real-world projects.

This work contributes to the development of advanced AI systems while offering flexible remote opportunities for people with software engineering experience.

The Role

As an LLM Evaluation & Repository Validation Engineer, you will help build verifiable software engineering tasks from public repository histories. You will analyze GitHub issues, configure code repositories, establish development environments, review test coverage and code quality, and run projects locally to assess bug-fixing performance.

You will also help identify repositories and issues that are especially challenging for large language models, contributing technical judgment from a software engineering perspective.

  • Contractor assignment
  • 20 hours per week
  • Three-month engagement
  • Remote worldwide
  • English-language work

What You'll Do

You will work with open-source software repositories and collaborate with researchers on repository selection and evaluation task design. The work requires hands-on technical investigation rather than purely theoretical review.

  • Analyze and triage GitHub issues across open-source libraries.
  • Set up and configure repositories, including Dockerization and development environments.
  • Evaluate unit test coverage and code quality.
  • Run, modify, and test real-world projects locally.
  • Assess LLM performance on software engineering and bug-fixing tasks.
  • Identify repositories and issues that present meaningful challenges for LLMs.
  • Collaborate with researchers on repository selection and task design.
  • Apply technical judgment from a senior software engineering perspective.

Requirements

This role is suited to someone who can confidently navigate complex codebases, debug software locally, and assess whether tests and fixes are reliable. Strong experience with public GitHub repositories is important.

  • Strong experience with at least one of Python, JavaScript, Java, Go, Rust, C, C++, C#, or Ruby.
  • Hands-on proficiency with Git, Docker, and basic software pipeline setup.
  • Ability to navigate complex codebases and debug locally.
  • Ability to run, modify, and test projects locally.
  • Experience evaluating unit tests, debugging code, and assessing bug-fixing performance.
  • Strong experience working with public GitHub repositories and complex codebases.

Helpful Background

The following experience is helpful for contributing effectively to repository validation and AI evaluation work, but is not presented as a mandatory requirement.

  • Tech lead-level software engineering experience.
  • Experience contributing to or evaluating open-source projects.
  • Familiarity with high-quality public GitHub repositories, including widely used repositories with 500 or more stars.
  • Previous work in LLM research or evaluation.
  • Experience with developer tools and automation agents.
  • Comfort collaborating with researchers on AI evaluation tasks.
  • Remote collaboration experience with overlapping Pacific Time hours.

Build Your AI Training Career with OpenTrain

AI training and data-labeling work spans code evaluation, model feedback, language, images, audio, and other forms of human-reviewed data. Contributors use their professional knowledge to help shape how modern AI systems perform.

Through OpenTrain, you can build a profile that showcases relevant experience, discover projects aligned with your skills, and develop a lasting portfolio in this rapidly growing field.

  • Work remotely with a flexible part-time schedule.
  • Apply software engineering expertise to cutting-edge AI evaluation.
  • Build credible experience in coding and LLM assessment.
  • Create a stronger profile for future AI training opportunities.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Rust LLM Evaluation and Repository Validation Engineer

Evaluate real Rust repositories and GitHub issues to help measure how LLMs solve software bugs. This part-time contractor role offers 20+ hours per week for engineers experienced with Rust, Git, Docker, and local testing.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

LLM Evaluation Software Engineer (Ruby)

Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Senior Software Engineer - C# LLM Evaluation & Code Validation

Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026