Rust LLM Evaluation and Repository Validation Engineer
Evaluate real Rust repositories and GitHub issues to help measure how LLMs solve software bugs. This part-time contractor role offers 20+ hours per week for engineers experienced with Rust, Git, Docker, and local testing.
Coding & Software
Remote
9 countries
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and grow in a rapidly expanding field.
Creating an OpenTrain account is free, and contributors can use the platform to find opportunities that match their technical experience.
About AI Training and Software Evaluation
AI training is the human work behind modern artificial intelligence. For software-focused projects, experienced engineers analyze code, create realistic programming tasks, and evaluate whether AI systems can understand, modify, and test real-world repositories.
This work places technical contributors close to the development of cutting-edge AI systems. Many AI-training projects are flexible and remote, making them suitable for people balancing project work with other commitments.
The Role
OpenTrain AI is seeking a Rust LLM Evaluation and Repository Validation Engineer to support training and evaluation datasets for realistic software engineering problems. You will work with open-source libraries, GitHub issues, repository environments, test suites, and local codebases to assess LLM performance in bug-fixing workflows.
This is an intermediate-level, part-time contractor opportunity requiring 20+ hours per week. The role is available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, and Mexico.
Employment type: Part-time contractor
Time commitment: 20+ hours per week
Working language: English
Primary focus: Rust software and LLM evaluation
What You'll Do
You will combine senior-level software engineering judgment with structured evaluation of repositories and coding tasks. Your work will help researchers identify challenging examples and assess how well LLMs handle realistic development scenarios.
Analyze and triage GitHub issues across open-source libraries.
Set up repositories through Dockerization, environment configuration, and basic pipeline setup.
Evaluate unit test coverage and test quality for software tasks.
Modify and run codebases locally to assess LLM performance in bug-fixing workflows.
Collaborate with researchers to identify repositories and issues that are challenging for LLMs.
Contribute technical judgment at a senior software engineering level, with potential opportunities to lead junior engineers on project work.
Required Skills and Background
The strongest candidates will have substantial Rust experience on real-world codebases and the ability to navigate complex repositories independently. Experience contributing to or evaluating open-source projects is helpful, as is prior work in LLM research or evaluation.
Strong Rust experience.
Strong Rust experience on real-world codebases.
Experience with Git, Docker, and repository setup workflows.
Ability to triage GitHub issues and assess unit test quality.
Ability to understand and navigate complex codebases.
Comfort running, modifying, debugging, and testing real-world projects locally.
Familiarity with open-source projects and local debugging.
Experience with developer tools or automation agents is helpful.
Prior LLM evaluation or research experience is a plus.
Senior-level software engineering experience, especially in repository analysis and testing, is helpful.
Why This Work Matters
Every major AI system depends on examples prepared and reviewed by people. By validating repositories, testing code, and judging software-engineering outputs, you help shape how AI systems reason about bugs and programming tasks.
This is an opportunity to apply practical Rust and software engineering expertise to a fast-growing area of technology while contributing to the development of more capable coding models.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Evaluate real software repositories and help measure how well AI coding systems fix bugs. This remote, three-month contractor assignment offers 20 hours per week for engineers skilled in codebases, testing, Git, and Docker.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.