Help improve AI-assisted software development by evaluating LLMs on real Go codebases, open-source issues, and bug-fixing tasks. Lead related projects while working remotely for 20+ hours per week.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 17, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps contributors discover specialized projects, build their AI-training careers, and apply in minutes.
Creating an OpenTrain account is free. Through this opportunity, you will contribute directly to the development of advanced AI systems while applying your software engineering expertise to practical evaluation work.
Remote contract opportunity
Part-time schedule of 20+ hours per week
Available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, and Mexico
English-language work
About AI Training and Code Evaluation
AI training is the human side of building artificial intelligence. Engineers and other specialists create, review, and evaluate examples that help AI models produce more useful and reliable results.
In this project, your software engineering judgment will help assess how large language models understand, modify, and troubleshoot real-world code. Your evaluations will support better AI-assisted software development.
Work on cutting-edge large language model evaluation
Use a human-in-the-loop approach to create verifiable software engineering tasks
Help identify the limits and capabilities of LLMs working with real code
The Role
OpenTrain is seeking a Senior LLM Code Evaluation Engineer to build LLM evaluation and training datasets from public GitHub repositories. This role combines practical software engineering with AI research and focuses on creating challenging, verifiable tasks for evaluating language models.
You will work with open-source libraries and real codebases, examining how effectively LLMs handle bugs, tests, repository setup, and software engineering workflows. You will also lead a team of junior engineers on related projects.
Minimum 3 years of overall software engineering experience
Strong proficiency in Go
20+ hours per week
Contractor and part-time engagement
What You'll Do
You will investigate public repositories and prepare realistic evaluation environments so that LLM performance can be assessed consistently and meaningfully. The work includes both hands-on coding and collaboration with researchers.
Analyze and triage GitHub issues across trending open-source libraries
Set up and configure code repositories
Dockerize projects and establish the required development environments
Evaluate unit test coverage and quality
Modify and run codebases locally to assess LLM performance in bug-fixing scenarios
Collaborate with researchers to identify repositories and issues that challenge LLMs
Lead a team of junior engineers on related projects
Required Qualifications
This role requires strong practical software engineering experience and the ability to work independently inside complex, real-world repositories. You should be comfortable configuring projects, changing code, and testing results locally.
At least 3 years of software engineering experience
Strong experience with the Go programming language
Proficiency with Git and Docker
Experience with basic software pipeline setup and environment automation
Ability to understand and navigate complex codebases
Experience analyzing and triaging GitHub issues in open-source projects
Skill in evaluating unit test coverage and quality
Comfort modifying, running, and testing real-world projects locally
Strong proficiency in Go
Helpful Background
The following experience is helpful for this opportunity, though it is not listed as required. It can help you contribute effectively to LLM-focused software engineering evaluation.
Previous participation in LLM research or evaluation projects
Experience building or testing developer tools
Experience working with automation agents
Why Work in AI Training
AI training and data-labeling work is a fast-growing part of the technology industry. Contributors help shape how modern AI systems behave by reviewing outputs, preparing training examples, and applying specialized expertise to challenging technical problems.
This opportunity offers flexible, remote work that can fit alongside other commitments while giving experienced engineers a direct role in improving state-of-the-art AI systems.
Remote work using a computer and internet connection
Flexible part-time participation
Direct impact on the future of AI-assisted software development
Opportunity to apply software engineering skills to emerging AI technology
How to Apply
Create a free OpenTrain account and apply through OpenTrain. Be prepared to demonstrate your Go expertise, software engineering experience, and ability to configure, modify, and evaluate real codebases.
Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.
Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.
Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.