Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect experienced contributors with hands-on projects that teach and shape how modern AI systems behave, and we hire and contract directly for this work.
About AI Training Work
AI training (data labeling and evaluation) is the human side of building today’s models: people prepare, test, and rate examples that models learn from. This role focuses on LLM evaluation and repository validation—practical, engineering-focused tasks that improve model behavior on real developer workflows.
The Role
You will be a senior contractor responsible for building and evaluating verifiable software engineering tasks drawn from public Git repositories, assessing how well LLMs handle bug fixing and developer scenarios. This is hands-on engineering work combined with structured evaluation for LLM training datasets.
Analyze and triage GitHub issues across trending open-source libraries.
Set up, Dockerize, and configure repositories and local environments to run tests and reproduce issues.
Evaluate unit test coverage, test quality, and identify gaps relevant to LLM evaluation.
Modify and run codebases locally to simulate LLM-driven bug fixes and measure outcomes.
Collaborate with researchers to select repositories and issues that challenge current LLMs.
Lead and mentor a small team of junior engineers on collaborative evaluation projects.
Requirements
Candidates must meet the technical and scheduling requirements below. The posting lists the experience level as Entry level in structured data, but the role explicitly requires a minimum of 3+ years of software engineering experience—please ensure you meet that requirement.
Minimum 3+ years of overall software engineering experience.
Strong experience with Ruby and comfort modifying Ruby codebases.
Proficiency with Git, Docker, and basic software pipeline setup.
Ability to understand and navigate complex, unfamiliar codebases.
Comfort running, modifying, and testing real-world projects locally.
Experience working in evaluation or LLM-related projects is required for some assignments.
Helpful Background
The following experiences will make you more effective in this role but are not strictly required.
Experience contributing to or evaluating open-source projects.
Previous participation in LLM research, model evaluation, or fine-tuning efforts.
Experience building or testing developer tools, automation agents, or CI workflows.
Prior experience leading or mentoring junior engineers on technical tasks.
Contractor Details and Scheduling
This is a remote contractor assignment through OpenTrain AI. You must commit to at least 20 hours per week and provide a 4-hour overlap with Pacific Standard Time (PST). Options of 20, 30, or 40 hours per week are available.
Allowed locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.
Employment types: Contractor, Part-time.
Language: English required for documentation and collaboration.
Work contributes to evaluation labels including EVALUATION_RATING, FINE_TUNING, and COMPUTER_PROGRAMMING_CODING.
How to Apply
Create or sign in to your OpenTrain account (free) to apply. Submit your resume/CV and examples of relevant Ruby or open-source work. Include notes on previous LLM evaluation experience, repository triage, or Dockerized environment setups where possible.
Applications should highlight Ruby projects, Git/Docker experience, and any open-source contributions.
Be prepared to demonstrate troubleshooting steps or reproduce a small issue during the evaluation process.
Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.