Evaluate how large language models solve realistic coding and bug-fixing tasks across open-source repositories. This worldwide, part-time contractor role offers hands-on AI training work for engineers comfortable with Git, Docker, and real codebases.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is recruiting contractors for specialized projects where technical professionals help improve the systems shaping the future of artificial intelligence.
Worldwide opportunity
Contractor and part-time engagement
20+ hours per week
Work conducted in English
About AI Training and LLM Evaluation
AI training is the human side of building artificial intelligence. People prepare examples, review model outputs, and evaluate performance so modern AI systems can become more accurate, useful, and reliable.
In this role, your software engineering judgment will help assess whether large language models can understand, modify, and repair real-world code. Your technical reviews will contribute to structured datasets and meaningful evaluations of coding capabilities.
Work directly with cutting-edge language model evaluation
Use realistic software engineering tasks rather than abstract coding exercises
Help shape how AI systems perform on development workflows
The Role
OpenTrain is seeking a software engineering contractor to evaluate large language models against realistic coding problems and validate repositories used to create AI training datasets. The work combines practical development, repository quality assessment, and structured evaluation of model performance.
You will work with public open-source codebases, investigate bug-fixing scenarios, and help produce verifiable software engineering tasks through hands-on technical review. The role is listed as entry level, while requiring practical programming and repository experience.
Role type: Contractor and part time
Time requirement: 20+ hours per week
Location: Worldwide
Language: English
Subject area: LLM evaluation and repository validation
What You’ll Do
You will examine real repositories and development tasks to determine whether they provide useful, technically sound challenges for language models. This includes setting up projects locally, reviewing tests, investigating issues, and evaluating model-generated fixes.
Analyze and triage issues across widely used open-source libraries
Configure repositories and development environments, including Docker-based setups
Review unit-test coverage and assess test quality
Modify and run codebases locally to evaluate LLM performance on bug-fixing tasks
Collaborate with researchers to identify repositories and issues that create meaningful challenges for language models
Help coordinate junior engineers on project work when needed
Required Skills
You should be able to work independently in complex software projects and make careful judgments about code quality, testing, repository setup, and model performance. Strong programming ability in at least one listed language is required.
Strong programming ability in at least one of Python, JavaScript, Java, Go, Rust, C, C++, C#, or Ruby
Practical proficiency with Git, Docker, and basic software pipeline setup
Ability to understand and navigate complex codebases
Experience running, modifying, and testing real-world projects locally
Ability to triage issues across open-source repositories
Judgment for assessing unit-test coverage and LLM bug-fixing performance
Helpful Background
Experience contributing to or evaluating open-source projects is useful, particularly with well-maintained public repositories. Previous work in LLM research or evaluation, or experience building and testing developer tools or automation agents, is also relevant.
Familiarity with widely used public repositories
Experience with repositories that have 500 or more stars
Background in LLM research or evaluation
Experience building or testing developer tools
Experience with automation agents
Why Build a Career in AI Training
AI training and data labeling are among the fastest-growing ways to work in technology. Contributors with programming expertise can apply their skills to projects that directly influence how state-of-the-art AI systems behave, often through flexible remote work.
Work remotely from anywhere with an internet connection
Choose flexible part-time work that can fit around other commitments
Build experience in a rapidly growing AI field
Develop a portfolio of specialized technical evaluation work
Apply Through OpenTrain
Create a free OpenTrain account to build your profile and apply in minutes. OpenTrain helps you discover AI training opportunities, present credible experience, and grow your work into a longer-term AI training portfolio.
Review the role and confirm your technical fit
Create or update your free OpenTrain profile
Apply through OpenTrain
Showcase your programming and repository experience
Help improve AI-assisted software development by evaluating LLMs on real Go codebases, open-source issues, and bug-fixing tasks. Lead related projects while working remotely for 20+ hours per week.
Evaluate how language models solve real software engineering problems using C#, GitHub repositories, Docker, and unit-test analysis. This remote contractor role requires 20 or more hours weekly and four hours of daily PST overlap.
Help improve large language models by curating code, evaluating AI-generated solutions, and building verification agents. This flexible contractor role supports 10 to 40 hours weekly for engineers in the US, Canada, and Western Europe.