LLM Evaluation Software Engineer Ruby
Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.
Posted Jul 20, 2026
Use your Rust expertise to validate repositories, assess test quality, and evaluate how LLMs solve real-world bugs. This flexible, 20+ hour remote contract is open to candidates in nine countries.
Coding & Software
9 countries
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open to applicants in
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a professional AI training profile, and apply in minutes. Creating an OpenTrain account is free.
AI training is the human side of building artificial intelligence. In software-focused projects, experienced engineers help create realistic coding datasets and evaluate whether AI systems can understand repositories, diagnose issues, and produce effective fixes.
This work gives software professionals a direct role in shaping how cutting-edge language models perform on real engineering tasks. Projects may involve reviewing code, testing solutions, and applying technical judgment to model outputs.
OpenTrain is seeking a Rust LLM Evaluation and Repository Validation Engineer to support training and evaluation datasets for realistic software engineering problems. You will analyze open-source repositories and GitHub issues, configure projects locally, assess test quality, and run code to evaluate how LLMs handle bug-fixing scenarios.
The role is suited to an intermediate-level engineer with strong Rust experience and the ability to navigate complex codebases. Senior-level technical judgment is important, with potential to lead junior engineers on project work.
You will work directly with repositories and software tasks that reflect real-world engineering challenges. Your assessments will help researchers identify difficult cases for LLMs and improve the quality of AI training and evaluation data.
You should be comfortable working with real-world Rust codebases, investigating software issues, and setting up projects for local execution. Experience contributing to or evaluating open-source projects is beneficial, as is prior work in LLM research or evaluation.
Every major AI system depends on human-reviewed examples and evaluations. By examining real repositories and challenging bug-fixing tasks, you will help make software-focused AI systems more capable, reliable, and useful to developers.
Create a free OpenTrain account, build your profile around your Rust and software engineering experience, and apply to this project in minutes. OpenTrain brings together opportunities in AI training so you can grow a lasting career in this rapidly expanding field.
Keep exploring
Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.
Posted Jul 20, 2026
Evaluate how well large language models solve real software bugs by configuring public repositories, testing code locally, and assessing unit tests. This remote, three-month contractor assignment offers 20 hours per week through OpenTrain.
Posted Jul 16, 2026
Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Posted Jul 17, 2026
Browse related job pages
Locations
Languages