Skip to content
OpenTrain AIFor AI Companies

Senior Python Software Engineer - LLM Evaluation

Help evaluate how AI models solve real-world software bugs by triaging GitHub issues, configuring repositories, and testing Python codebases. This part-time contractor role requires 20+ hours per week and is open in select countries.

OpenTrain AI

Coding & Software

Remote

9 countries

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open to applicants in

India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build an AI training profile, and apply in minutes.

Creating an OpenTrain account is free, and this opportunity is available as part-time contract work for contributors in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.

About AI Training and LLM Evaluation

AI training is the human side of building artificial intelligence. Engineers and other specialists prepare examples, test model outputs, and provide evaluations that help modern AI systems perform better.

In this project, your software engineering expertise will help assess how well large language models handle realistic debugging and bug-fixing tasks across open-source repositories.

The Role

OpenTrain AI is seeking a Senior Software Engineer specializing in Python, LLM evaluation, and repository validation. You will work with realistic software engineering problems by analyzing GitHub issues, configuring development environments, evaluating tests, and running code locally to assess model performance.

The work may also involve helping researchers identify challenging repositories and issues and supporting junior engineers during collaborative project work.

  • Contractor and part-time engagement
  • Time requirement: 20+ hours per week
  • Working language: English
  • Listed experience level: Entry level

What You'll Do

You will prepare and assess real-world software projects so that LLM performance can be evaluated against meaningful engineering tasks. The work requires hands-on interaction with codebases, repositories, development environments, and test suites.

  • Analyze and triage GitHub issues across open-source libraries.
  • Set up repositories using Docker and development environment automation.
  • Evaluate unit test coverage and test quality.
  • Run and modify codebases locally to assess model performance on bug-fixing tasks.
  • Collaborate with researchers on repository and issue selection.
  • Support junior engineers on collaborative project work when needed.

Required Skills and Experience

This role calls for at least three years of overall engineering experience along with strong practical Python skills. You should be comfortable working directly in complex codebases, setting up repositories, and testing real-world projects locally.

  • 3+ years of overall engineering experience.
  • Strong Python experience, including software engineering and code modification.
  • Hands-on skill with Git, Docker, and software pipeline setup.
  • Ability to triage GitHub issues and evaluate unit test quality.
  • Comfort running, modifying, and debugging real-world code locally.
  • Familiarity with open-source repositories and high-quality public GitHub projects.

Helpful Background

Experience contributing to or evaluating open-source projects is valuable. Background in LLM research or evaluation can also help, particularly when assessing whether a model has correctly addressed a software engineering issue.

  • Experience building or testing developer tools.
  • Experience with automation agents.
  • Prior LLM research or evaluation experience.
  • Familiarity with complex, high-quality public GitHub repositories.

Why Work in AI Training

AI training and data-labeling work is a fast-growing part of the technology industry. Contributors help shape how state-of-the-art AI systems behave by reviewing outputs, evaluating performance, and preparing high-quality training data.

Many projects offer flexible, part-time schedules and can be completed remotely with a computer and internet connection. This engineering-focused opportunity lets you apply your development skills to cutting-edge model evaluation.

  • Contribute directly to the development of AI systems.
  • Apply software engineering skills to realistic coding and debugging tasks.
  • Build experience in an expanding AI training field.
  • Work on a part-time schedule of 20+ hours per week.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Senior Python Software Engineer – LLM Evaluation

Help evaluate how large language models work with real Python code. Build repository environments, triage issues, assess tests, and analyze LLM bug-fixing performance in a flexible, remote contract role.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

LLM Evaluation Software Engineer (Ruby)

Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Senior LLM Code Evaluation Engineer

Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026