Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 17, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the #1 platform for people building careers in AI training and data labeling. We connect skilled contributors with meaningful projects that shape how state-of-the-art AI systems behave. OpenTrain AI is the hiring and contracting organization for this role.
Why AI training matters
AI training (data labeling and human feedback work) is the human side of building intelligent systems: people create, review, and rate examples that modern models learn from. This work is often remote, flexible, and accessible — contributors directly influence how developer-facing AI tools perform.
Role overview
We're hiring a Senior LLM Code Evaluation Engineer to build evaluation and training datasets from public open-source code repositories. This role blends practical software engineering with LLM evaluation research: you will identify challenging code, prepare reproducible testbeds, run and modify projects locally, and lead junior engineers on related tasks.
Part-time contractor role — minimum commitment: 20+ hours per week.
Work language: English.
Candidate eligibility restricted to: IN, PK, NG, KE, EG, GH, BD, TR, MX.
What you'll do
You will work hands-on with real codebases to create verifiable engineering tasks that test LLM capabilities in debugging, code generation, and repair. Expect a mix of independent technical work and collaboration with researchers and reviewers.
Analyze and triage issues across trending open-source libraries to find high-value evaluation cases.
Set up and configure repositories, including Dockerization and environment automation.
Assess unit test coverage and the quality of existing tests.
Modify and run codebases locally to reproduce bugs and evaluate LLM-assisted fixes.
Collaborate with researchers to select repositories and design evaluation scenarios.
Lead and mentor a small team of junior engineers on evaluation tasks and quality standards.
Required qualifications
Your application should demonstrate practical software engineering experience focused on working with real-world codebases and automation.
Minimum 3 years of overall software engineering experience.
Strong proficiency in Go (Golang).
Proficient with Git and Docker; able to create environment automation and containerized setups.
Comfortable understanding, navigating, modifying, and running complex codebases locally.
Experience evaluating unit tests and judging test quality.
Helpful experience
The following backgrounds are useful but not strictly required; they will help you hit the ground running and influence dataset design and evaluation rigor.
Participation in LLM research, evaluation, or annotation projects.
Experience building or testing developer tools, CI automation, or agent-style workflows.
Prior mentoring or team lead experience on engineering projects.
Who should apply
Apply if you enjoy digging into messy, real-world code; designing reproducible tests; and translating engineering problems into clear evaluation labels. This role suits engineers who want part-time, impactful work shaping how developer-facing LLMs perform.
Open to contractors seeking flexible, remote part-time work (20+ hrs/week).
Fluent English required for documentation and collaboration.
Candidates must be located in one of the eligible countries listed in the role overview.
How the work is organized
Tasks combine hands-on repository setup, annotation/evaluation labeling, and collaboration with researchers. You'll produce reproducible tasks, label outcomes (coding and evaluation ratings), and lead junior contributors through review cycles. OpenTrain AI provides the project scope and coordination.
Label types include computer programming/coding and evaluation/rating.
Expect a human-in-the-loop workflow: set up a repo, run tests, generate evaluation prompts, and record structured ratings.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.
Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.