Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people start and grow careers teaching AI. We help contractors discover projects, build a durable portfolio of AI training work, and manage opportunities in one place so proof of work and skills follow you across roles.
For this engagement, OpenTrain is the hiring and contracting organization. You'll join a distributed team contributing directly to how modern AI systems learn from code and developer workflows.
Why AI Training Work Matters (and What You'll Be Doing)
AI training — also called data labeling or evaluation — is the human side of building intelligent systems. For coding-focused projects, contributors create and validate examples that teach models to understand, run, and fix real software.
In this role you'll shape LLM behavior on software engineering tasks by selecting repositories and issues, preparing runnable codebases, and measuring how well models perform at debugging and repair.
Role Overview — What You'll Do
As a contractor Senior Software Engineer focused on LLM evaluation and repository validation, you'll work hands-on with open-source codebases to create robust evaluation tasks and validate model outputs in realistic development environments.
Analyze and triage issues across trending open-source libraries to identify high-quality evaluation candidates.
Set up and configure repositories locally, including Dockerization and environment setup, so tasks run reliably.
Evaluate unit test coverage and test quality to determine whether a repo/issue is suitable for LLM evaluation.
Modify and run codebases locally to reproduce bugs and assess LLM-generated fixes in practice.
Collaborate with researchers to design challenging evaluation scenarios and select repositories that test model limits.
Opportunity to lead or mentor a small team of junior engineers on validation tasks.
Requirements
We need engineers who can move quickly in real code and reliably produce reproducible validation artifacts. The list below captures required and desired skills — all are derived from the role description.
Minimum 3+ years of overall software engineering experience.
Strong experience with C# development and ecosystems.
Proficiency with Git, Docker, and basic pipeline/setup for running projects locally.
Comfort with navigating complex codebases, running tests, and making code changes to reproduce bugs.
Experience contributing to or evaluating open-source projects is a plus.
Previous participation in LLM research or evaluation projects is preferred but not required.
Commitment, Location & Logistics
This is a fully remote contractor assignment. You should expect at least 20 hours per week, with options for 30 or 40 hours. The role requires approximately 4 hours of overlap with Pacific Time (PST).
Open to applicants located in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico. English fluency is required for documentation and collaboration.
This is a contractor position and does not provide medical or paid leave benefits.
How the Work Is Measured and Tools
Tasks are evaluation and rating work: you'll produce reproducible validation steps, run tests, and score model outputs against deterministic criteria. Deliverables typically include runnable repositories, test results, issue triage notes, and evaluation ratings.
You should be comfortable using standard developer tools (IDE, Git, Docker, test runners) and documenting setup and reproduction steps so other evaluators can replicate your work.
Who Should Apply
Apply if you enjoy hands-on engineering work, are curious about how LLMs interact with real code, and want to help shape model behavior for debugging and repair scenarios. This role fits developers who like reproducible workflows, testing, and collaborating with research teams.
Junior engineers with strong C# skills and the ability to run and validate projects locally may also be considered for mentorship-track positions.
Next Steps — How to Apply
Create or update your OpenTrain profile, include samples or descriptions of C# and open-source work, and apply to this role. Provide concise examples of repositories you’ve worked on and any prior experience with evaluation or research-style tasks.
If selected, you will be given onboarding instructions, coding and validation tasks, and collaboration guidelines to begin contributing.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.