Help evaluate how large language models work with real Python code. Build repository environments, triage issues, assess tests, and analyze LLM bug-fixing performance in a flexible, remote contract role.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the platform where people discover and build careers in AI training and data labeling. We connect contributors with specialized projects, help them build an AI training profile, and make it possible to apply in minutes.
Remote contract opportunity
Part-time commitment of 20+ hours per week
Open to applicants in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Türkiye, and Mexico
About AI Training and LLM Evaluation
AI training is the human side of building modern artificial intelligence. Engineers, researchers, and other specialists prepare examples and evaluate model behavior so AI systems can become more capable, reliable, and useful.
In this role, your software engineering expertise will help assess how large language models interact with real-world code. Your evaluations can influence the development of AI-assisted software tools and the datasets used to improve them.
The Role
OpenTrain AI is recruiting a Senior Software Engineer specializing in Python to support LLM evaluation and repository validation. You will work hands-on with open-source codebases, development environments, issue triage, test coverage, and software quality.
The work is suited to an experienced engineer who can navigate complex repositories, run and modify projects locally, and assess whether code and tests provide a meaningful challenge for an AI system.
Focus area: LLM software engineering evaluation
Core language: Python
Work arrangement: Part-time contractor
Time requirement: 20+ hours per week
Working language: English
What You'll Do
You will help create and evaluate realistic software engineering tasks for large language models. This includes preparing repositories, understanding issues, running projects, and judging the quality of tests and potential solutions.
Analyze and triage GitHub issues across trending open-source libraries.
Set up and configure code repositories, including Dockerization and development environment setup.
Evaluate unit test coverage and overall test quality.
Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.
Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.
Potentially lead a team of junior engineers.
Requirements
You should be comfortable working independently in real-world software repositories and evaluating both implementation quality and testing practices. Strong Python experience is essential, along with practical familiarity with the tools used to set up and run codebases.
At least 3 years of overall professional experience.
Strong experience with Python.
Proficiency with Git, Docker, and basic software pipeline setup.
Ability to understand and navigate complex codebases.
Comfort running, modifying, and testing real-world projects locally.
Experience contributing to or evaluating open-source projects.
Fluency in English.
Helpful Background
The following experience can help you succeed in this project, particularly when evaluating difficult repository-level tasks and understanding how AI systems approach software engineering problems.
Participation in LLM research or evaluation projects.
Experience building or testing developer tools.
Experience working with automation agents.
Background collaborating with researchers or engineering teams.
Build the Future of AI-Assisted Development
This is an opportunity to apply your Python and software engineering judgment to cutting-edge AI work. By examining real repositories, test suites, and bug-fixing behavior, you will help shape how future models support developers.
Create a free OpenTrain account to build your profile and apply for AI training opportunities in minutes.
Help evaluate how AI models solve real-world software bugs by triaging GitHub issues, configuring repositories, and testing Python codebases. This part-time contractor role requires 20+ hours per week and is open in select countries.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.