Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.
About AI Training and LLM Evaluation
AI training is the human side of building modern artificial intelligence. Engineers and other specialists prepare examples, evaluate model outputs, and test how well AI systems perform on realistic tasks. In this role, your software engineering judgment will help assess and improve large language models, or LLMs, working with real open-source code and bug-fixing scenarios.
Contribute to cutting-edge LLM training and evaluation work
Work remotely with flexible part-time or higher-hour options
Help shape how AI systems handle software development tasks
The Role
OpenTrain is recruiting a software engineer contractor focused on LLM evaluation and repository validation. You will build and evaluate verifiable software engineering tasks from public GitHub repositories, contributing to LLM training datasets. The work involves analyzing open-source code, triaging issues, and assessing model performance on real-world bug-fixing and software development scenarios.
Role focus: LLM evaluation, repository validation, and software engineering tasks
Primary language: Ruby
Work arrangement: Remote contractor assignment
Availability: 20, 30, or 40 hours per week
Minimum commitment: 20 hours per week
Schedule requirement: 4-hour overlap with PST
What You'll Do
You will work directly with codebases and repository environments to create reliable evaluation tasks and judge how LLMs perform. The role also includes collaboration and technical leadership on team projects.
Analyze and triage GitHub issues across trending open-source libraries
Set up and configure code repositories, including Dockerization and environment setup
Evaluate unit test coverage and quality
Modify and run codebases locally to assess LLM performance in bug-fixing scenarios
Collaborate with researchers to design and identify repositories and issues that challenge LLMs
Lead a team of junior engineers on collaborative projects
Required Skills and Experience
This assignment requires practical software engineering experience and the ability to navigate complex repositories independently. You should be comfortable running, modifying, and testing real-world projects locally while using Ruby, Git, and Docker.
At least 3 years of overall software engineering experience
Strong experience with Ruby
Proficiency with Git and Docker
Basic software pipeline setup skills
Ability to understand and navigate complex codebases
Comfort running, modifying, and testing real-world projects locally
Experience setting up and configuring code repositories
Skill analyzing and triaging GitHub issues
Experience evaluating or testing codebases and unit test coverage
Comfort working with LLM evaluation or AI training datasets
Helpful Background
The following experience can help you contribute effectively to repository-based LLM evaluation projects, though it is listed as helpful background rather than a required qualification.
Contributing to or evaluating open-source projects
Participating in LLM research or evaluation projects
Building or testing developer tools
Building or testing automation agents
Location and Contractor Details
This is a remote contractor opportunity available only to candidates located in the listed countries. The role requires at least 20 hours each week and a 4-hour overlap with Pacific Standard Time.
Eligible locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico
Weekly options: 20, 30, or 40 hours
Required availability: 20 or more hours per week
Required schedule overlap: 4 hours with PST
Language: English
How to Apply Through OpenTrain
Create a free OpenTrain account, build your profile around your software engineering and Ruby experience, and apply in minutes. OpenTrain brings together opportunities in AI training and data labeling so you can develop a durable career working on how advanced AI systems are built and evaluated.
Highlight Ruby, Git, Docker, and repository-testing experience
Include relevant open-source, LLM evaluation, or developer-tool work
Confirm your location and weekly availability before applying
Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.
Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.