Skip to content
OpenTrain AIFor AI Companies

Rust LLM Evaluation Engineer

Use your Rust expertise to validate repositories, assess test quality, and evaluate how LLMs solve real-world bugs. This flexible, 20+ hour remote contract is open to candidates in nine countries.

OpenTrain AI

Coding & Software

Remote

9 countries

Eligibility

Intermediate

Experience

Jul 16, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.

  • Flexible contractor and part-time opportunity
  • Remote work for eligible candidates in the listed countries
  • 20+ hours per week

About AI Training and Software Evaluation

AI training is the human side of building artificial intelligence. In software-focused projects, experienced engineers help create realistic coding datasets and evaluate whether AI systems can understand repositories, diagnose issues, and produce effective fixes.

This work gives software professionals a direct role in shaping how cutting-edge language models perform on real engineering tasks. Projects may involve reviewing code, testing solutions, and applying technical judgment to model outputs.

  • Contribute to datasets for realistic software engineering problems
  • Evaluate LLM performance on bug-fixing workflows
  • Apply practical engineering judgment to emerging AI systems

The Role

OpenTrain is seeking a Rust LLM Evaluation and Repository Validation Engineer to support training and evaluation datasets for realistic software engineering problems. You will analyze open-source repositories and GitHub issues, configure projects locally, assess test quality, and run code to evaluate how LLMs handle bug-fixing scenarios.

The role is suited to an intermediate-level engineer with strong Rust experience and the ability to navigate complex codebases. Senior-level technical judgment is important, with potential to lead junior engineers on project work.

  • Role type: Contractor and part-time
  • Time requirement: 20+ hours per week
  • Working language: English
  • Experience level: Intermediate
  • Eligible countries: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, and Mexico

What You’ll Do

You will work directly with repositories and software tasks that reflect real-world engineering challenges. Your assessments will help researchers identify difficult cases for LLMs and improve the quality of AI training and evaluation data.

  • Analyze and triage GitHub issues across open-source libraries.
  • Set up repositories, including Dockerization and environment configuration.
  • Evaluate unit test coverage and the quality of tests for software tasks.
  • Modify and run codebases locally to assess LLM performance in bug-fixing workflows.
  • Collaborate with researchers to identify repositories and issues that challenge LLMs.
  • Provide technical judgment at a senior software engineering level.
  • Potentially lead junior engineers on project work.

Requirements and Helpful Experience

You should be comfortable working with real-world Rust codebases, investigating software issues, and setting up projects for local execution. Experience contributing to or evaluating open-source projects is beneficial, as is prior work in LLM research or evaluation.

  • Strong Rust experience on real-world codebases
  • Ability to understand and navigate complex codebases
  • Experience triaging GitHub issues
  • Ability to assess unit test coverage and quality
  • Experience with Git, Docker, and basic software pipeline setup
  • Comfort running, modifying, and testing real-world projects locally
  • Open-source contribution or evaluation experience is a plus
  • Prior LLM research or evaluation experience is helpful
  • Experience with developer tools or automation agents is helpful
  • Senior-level software engineering experience in repository analysis and testing is helpful

Why This Work Matters

Every major AI system depends on human-reviewed examples and evaluations. By examining real repositories and challenging bug-fixing tasks, you will help make software-focused AI systems more capable, reliable, and useful to developers.

  • Work on cutting-edge AI training and evaluation
  • Use specialized Rust and software engineering expertise
  • Choose flexible work that can fit around other commitments
  • Help shape how AI systems handle practical coding problems

How to Apply Through OpenTrain

Create a free OpenTrain account, build your profile around your Rust and software engineering experience, and apply to this project in minutes. OpenTrain brings together opportunities in AI training so you can grow a lasting career in this rapidly expanding field.

  • Highlight Rust, Git, Docker, testing, and repository analysis experience.
  • Mention open-source, debugging, automation-agent, or LLM evaluation work.
  • Confirm that you can commit to 20+ hours per week.
  • Apply through OpenTrain if you are located in an eligible country.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

LLM Evaluation Software Engineer Ruby

Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

LLM Evaluation and Repository Validation Engineer

Evaluate how well large language models solve real software bugs by configuring public repositories, testing code locally, and assessing unit tests. This remote, three-month contractor assignment offers 20 hours per week through OpenTrain.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

C++ LLM Evaluation Software Engineer

Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026