Skip to content
OpenTrain AIFor AI Companies

LLM Evaluation Software Engineer Ruby

Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.

OpenTrain AI

Coding & Software

Remote

9 countries

Eligibility

Entry

Experience

Jul 20, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.

About AI Training and LLM Evaluation

AI training is the human side of building modern artificial intelligence. Engineers and other specialists prepare examples, evaluate model outputs, and test how well AI systems perform on realistic tasks. In this role, your software engineering judgment will help assess and improve large language models, or LLMs, working with real open-source code and bug-fixing scenarios.

  • Contribute to cutting-edge LLM training and evaluation work
  • Work remotely with flexible part-time or higher-hour options
  • Help shape how AI systems handle software development tasks

The Role

OpenTrain is recruiting a software engineer contractor focused on LLM evaluation and repository validation. You will build and evaluate verifiable software engineering tasks from public GitHub repositories, contributing to LLM training datasets. The work involves analyzing open-source code, triaging issues, and assessing model performance on real-world bug-fixing and software development scenarios.

  • Role focus: LLM evaluation, repository validation, and software engineering tasks
  • Primary language: Ruby
  • Work arrangement: Remote contractor assignment
  • Availability: 20, 30, or 40 hours per week
  • Minimum commitment: 20 hours per week
  • Schedule requirement: 4-hour overlap with PST

What You'll Do

You will work directly with codebases and repository environments to create reliable evaluation tasks and judge how LLMs perform. The role also includes collaboration and technical leadership on team projects.

  • Analyze and triage GitHub issues across trending open-source libraries
  • Set up and configure code repositories, including Dockerization and environment setup
  • Evaluate unit test coverage and quality
  • Modify and run codebases locally to assess LLM performance in bug-fixing scenarios
  • Collaborate with researchers to design and identify repositories and issues that challenge LLMs
  • Lead a team of junior engineers on collaborative projects

Required Skills and Experience

This assignment requires practical software engineering experience and the ability to navigate complex repositories independently. You should be comfortable running, modifying, and testing real-world projects locally while using Ruby, Git, and Docker.

  • At least 3 years of overall software engineering experience
  • Strong experience with Ruby
  • Proficiency with Git and Docker
  • Basic software pipeline setup skills
  • Ability to understand and navigate complex codebases
  • Comfort running, modifying, and testing real-world projects locally
  • Experience setting up and configuring code repositories
  • Skill analyzing and triaging GitHub issues
  • Experience evaluating or testing codebases and unit test coverage
  • Comfort working with LLM evaluation or AI training datasets

Helpful Background

The following experience can help you contribute effectively to repository-based LLM evaluation projects, though it is listed as helpful background rather than a required qualification.

  • Contributing to or evaluating open-source projects
  • Participating in LLM research or evaluation projects
  • Building or testing developer tools
  • Building or testing automation agents

Location and Contractor Details

This is a remote contractor opportunity available only to candidates located in the listed countries. The role requires at least 20 hours each week and a 4-hour overlap with Pacific Standard Time.

  • Eligible locations: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico
  • Weekly options: 20, 30, or 40 hours
  • Required availability: 20 or more hours per week
  • Required schedule overlap: 4 hours with PST
  • Language: English

How to Apply Through OpenTrain

Create a free OpenTrain account, build your profile around your software engineering and Ruby experience, and apply in minutes. OpenTrain brings together opportunities in AI training and data labeling so you can develop a durable career working on how advanced AI systems are built and evaluated.

  • Highlight Ruby, Git, Docker, and repository-testing experience
  • Include relevant open-source, LLM evaluation, or developer-tool work
  • Confirm your location and weekly availability before applying

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Senior Software Engineer LLM Evaluation

Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

C++ LLM Evaluation Software Engineer

Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

Senior C++ Software Engineer for LLM Evaluation

Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026