Skip to content
OpenTrain AIFor AI Companies

Senior LLM Code Evaluation Engineer

Help improve AI-assisted software development by evaluating LLMs on real Go codebases, open-source issues, and bug-fixing tasks. Lead related projects while working remotely for 20+ hours per week.

OpenTrain AI

Coding & Software

Remote

9 countries

Eligibility

Entry

Experience

Jul 17, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps contributors discover specialized projects, build their AI-training careers, and apply in minutes.

Creating an OpenTrain account is free. Through this opportunity, you will contribute directly to the development of advanced AI systems while applying your software engineering expertise to practical evaluation work.

  • Remote contract opportunity
  • Part-time schedule of 20+ hours per week
  • Available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, and Mexico
  • English-language work

About AI Training and Code Evaluation

AI training is the human side of building artificial intelligence. Engineers and other specialists create, review, and evaluate examples that help AI models produce more useful and reliable results.

In this project, your software engineering judgment will help assess how large language models understand, modify, and troubleshoot real-world code. Your evaluations will support better AI-assisted software development.

  • Work on cutting-edge large language model evaluation
  • Use a human-in-the-loop approach to create verifiable software engineering tasks
  • Help identify the limits and capabilities of LLMs working with real code

The Role

OpenTrain is seeking a Senior LLM Code Evaluation Engineer to build LLM evaluation and training datasets from public GitHub repositories. This role combines practical software engineering with AI research and focuses on creating challenging, verifiable tasks for evaluating language models.

You will work with open-source libraries and real codebases, examining how effectively LLMs handle bugs, tests, repository setup, and software engineering workflows. You will also lead a team of junior engineers on related projects.

  • Minimum 3 years of overall software engineering experience
  • Strong proficiency in Go
  • 20+ hours per week
  • Contractor and part-time engagement

What You'll Do

You will investigate public repositories and prepare realistic evaluation environments so that LLM performance can be assessed consistently and meaningfully. The work includes both hands-on coding and collaboration with researchers.

  • Analyze and triage GitHub issues across trending open-source libraries
  • Set up and configure code repositories
  • Dockerize projects and establish the required development environments
  • Evaluate unit test coverage and quality
  • Modify and run codebases locally to assess LLM performance in bug-fixing scenarios
  • Collaborate with researchers to identify repositories and issues that challenge LLMs
  • Lead a team of junior engineers on related projects

Required Qualifications

This role requires strong practical software engineering experience and the ability to work independently inside complex, real-world repositories. You should be comfortable configuring projects, changing code, and testing results locally.

  • At least 3 years of software engineering experience
  • Strong experience with the Go programming language
  • Proficiency with Git and Docker
  • Experience with basic software pipeline setup and environment automation
  • Ability to understand and navigate complex codebases
  • Experience analyzing and triaging GitHub issues in open-source projects
  • Skill in evaluating unit test coverage and quality
  • Comfort modifying, running, and testing real-world projects locally
  • Strong proficiency in Go

Helpful Background

The following experience is helpful for this opportunity, though it is not listed as required. It can help you contribute effectively to LLM-focused software engineering evaluation.

  • Previous participation in LLM research or evaluation projects
  • Experience building or testing developer tools
  • Experience working with automation agents

Why Work in AI Training

AI training and data-labeling work is a fast-growing part of the technology industry. Contributors help shape how modern AI systems behave by reviewing outputs, preparing training examples, and applying specialized expertise to challenging technical problems.

This opportunity offers flexible, remote work that can fit alongside other commitments while giving experienced engineers a direct role in improving state-of-the-art AI systems.

  • Remote work using a computer and internet connection
  • Flexible part-time participation
  • Direct impact on the future of AI-assisted software development
  • Opportunity to apply software engineering skills to emerging AI technology

How to Apply

Create a free OpenTrain account and apply through OpenTrain. Be prepared to demonstrate your Go expertise, software engineering experience, and ability to configure, modify, and evaluate real codebases.

  • Apply through OpenTrain
  • Confirm availability for 20+ hours per week
  • Review the location eligibility before applying
  • Highlight relevant Go, Git, Docker, GitHub, testing, and LLM evaluation experience

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Senior Software Engineer LLM Evaluation

Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

Senior Software Engineer, LLM Evaluation

Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 6, 2026

Senior C++ Software Engineer for LLM Evaluation

Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026