Skip to content
OpenTrain AIFor AI Companies

Python Software Engineering LLM Evaluator

Help evaluate how language models solve real Python software engineering and bug-fixing tasks. This fully remote, three-month contractor assignment offers $100 per task and requires 20+ hours weekly from eligible countries.

Apply now
OpenTrain AI

Coding & Software

Remote Per task · $100/label

$100/label

Compensation

8 countries

Eligibility

Entry

Experience

Sep 21, 2026

Posted

Open to applicants in

India
+ more
  • Bangladesh
  • Egypt
  • Ghana
  • India
  • Mexico
  • Nigeria
  • Pakistan
  • Türkiye

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover projects, build a professional profile, and apply to opportunities that match their skills. Creating an OpenTrain account is free.

  • Build a lasting portfolio of AI training and evaluation experience.
  • Find flexible contractor work connected to the rapidly growing AI industry.

About AI Training and LLM Evaluation

AI training is the human side of building artificial intelligence. For language models, skilled contributors review outputs, design realistic tasks, and assess whether systems can produce accurate, useful code. Your software engineering judgment can help improve how models handle real development challenges.

  • Work with realistic programming and bug-fixing scenarios.
  • Evaluate model behavior against real codebases and software engineering expectations.
  • Contribute to cutting-edge AI development through practical technical review.

The Role

OpenTrain is recruiting Python software engineers for a three-month contractor assignment focused on LLM evaluation and training. You will work with public GitHub repositories and realistic software engineering tasks to assess how language models handle real code and bug-fixing scenarios.

This is a fully remote, part-time assignment requiring 20+ hours per week. The role is available to candidates based in India, Pakistan, Nigeria, Egypt, Ghana, Bangladesh, Turkey, or Mexico, and pays $100 per task.

  • Role type: Remote contractor, part time
  • Assignment length: Three months
  • Time requirement: 20+ hours per week
  • Payment: $100 per task
  • Working language: English

What You'll Do

  • Analyze and triage GitHub issues across open-source libraries.
  • Configure repositories and Docker-based development environments.
  • Run, modify, and test real codebases locally.
  • Evaluate unit-test coverage and test quality.
  • Assess LLM performance on software engineering and bug-fixing tasks.
  • Collaborate with researchers to identify repositories and issues that provide meaningful challenges for LLMs.

Requirements

The structured role level is entry level, while the assignment specifically requires at least three years of software engineering experience. Candidates should be comfortable working independently with complex public codebases and evaluating both software quality and LLM performance.

  • At least three years of software engineering experience.
  • Strong Python software engineering experience.
  • Proficiency with Git, Docker, and basic software pipeline setup.
  • Ability to understand and navigate complex codebases.
  • Experience running, modifying, and testing real-world projects locally.
  • Judgment in evaluating unit-test coverage, test quality, and LLM bug-fixing performance.
  • English proficiency for technical evaluation work.

Preferred Experience

  • Experience contributing to or evaluating open-source projects.
  • Previous LLM research or evaluation experience.

Why Build an AI Training Career With OpenTrain

AI training and data-labeling work gives people a way to contribute directly to state-of-the-art AI systems. Many projects are remote and flexible, while specialized technical assignments let experienced professionals apply their existing expertise to emerging model-development work.

Through OpenTrain, you can build a profile that showcases your AI training experience, discover projects aligned with your skills, and develop a longer-term portfolio in this fast-growing field.

  • Work remotely with a flexible weekly commitment.
  • Apply software engineering expertise to advanced AI evaluation.
  • Grow experience in LLM testing, code assessment, and AI training.

How to Apply

Create or use your free OpenTrain account, review the assignment details, and submit your application. Make sure your profile reflects your Python, Git, Docker, software engineering, and code evaluation experience.

  • Confirm that you are based in an eligible country.
  • Highlight experience with complex codebases and real-world testing.
  • Apply through OpenTrain in minutes.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Senior Python Software Engineer LLM Evaluation

Evaluate how AI models fix real software bugs in open-source Python repositories. This flexible, part-time contractor role focuses on GitHub issue triage, Docker environments, testing, and LLM performance assessment.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

C++ LLM Evaluation Software Engineer

Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

LLM Evaluation Software Engineer Ruby

Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026