Skip to content
OpenTrain AIFor AI Companies

Senior Python Developer for AI Model Evaluation

Use advanced Python skills to evaluate AI models, create training data, rank responses, and improve coding-focused systems. This worldwide contract role offers flexible part-time work of 20+ hours per week.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Intermediate

Experience

Jul 16, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the organization hiring and contracting for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and this role is open to qualified contractors worldwide.

About AI Model Evaluation Work

AI training is the human work behind modern artificial intelligence. Contributors write prompts, create and review model responses, evaluate technical accuracy, and provide structured feedback that helps AI systems become more useful and reliable.

In this project, your Python expertise will support coding-focused evaluation, supervised fine-tuning data, and RLHF-style feedback workflows.

The Role

OpenTrain AI is seeking a Senior Python Developer for AI model evaluation and data generation work. You will combine software development experience with careful judgment to assess multiple large language model outputs and create high-quality examples for training and evaluation.

This is an intermediate-level, part-time contractor role requiring 20+ hours per week. The work is worldwide and conducted in English.

  • Employment type: Contractor
  • Schedule: Part time, 20+ hours per week
  • Location: Worldwide
  • Working language: English

What You'll Do

You will write, review, and explain code-based solutions while helping develop evaluation methods for AI models. Clear reasoning and consistent application of evaluation criteria are central to the work.

  • Write efficient Python code for AI training and evaluation workflows
  • Build responses and solutions for prompt-based coding tasks
  • Evaluate and rank model outputs against predefined criteria
  • Produce rationales and peer reviews that support model improvement
  • Contribute to supervised fine-tuning datasets and RLHF-style feedback workflows
  • Review code and documentation for quality, correctness, and regressions
  • Work with researchers, annotators, and cross-functional teams on model performance

Required Skills and Experience

You should have strong Python development experience and be comfortable assessing both code quality and AI-generated responses. The role requires clear written and conversational English communication.

  • 3+ years of strong Python programming experience
  • Experience with unit, integration, and property-based testing
  • Knowledge of multithreading and asynchronous programming in Python
  • Ability to refactor code safely and debug memory or concurrency issues
  • Familiarity with code quality, formatting, and software development best practices
  • Experience writing code for AI training or evaluation workflows
  • Ability to compare model outputs and explain evaluation decisions
  • Familiarity with SFT, RLHF, and benchmark-style model assessment
  • Fluent written and conversational English

Why This Work Matters

Every major AI system depends on people who prepare examples, review outputs, and identify where models need to improve. By combining your Python skills with thoughtful evaluation, you can contribute directly to how advanced AI systems behave while working remotely in a growing technical field.

  • Apply software engineering expertise to cutting-edge AI development
  • Help improve model alignment, clarity, and technical accuracy
  • Work remotely with a flexible part-time contractor schedule
  • Build experience in AI training, evaluation, and data generation

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Python AI Model Training & Evaluation Engineer

Build and evaluate AI systems through Python development, supervised fine-tuning, RLHF, response ranking, and public-data analysis. This fully remote contractor role offers 20, 30, or 40 hours per week for a one-month engagement.

Coding & Software
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

Python Backend Developer – AI Coding Evaluation

Build scalable Python APIs and help test AI coding tools in focused four-day bursts. This remote contractor role offers $50–$100 per hour, 20+ hours weekly, and the chance to shape developer-focused AI systems.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$100/hr

Posted Jun 28, 2026

Senior Python Software Engineer – LLM Evaluation

Help evaluate how large language models work with real Python code. Build repository environments, triage issues, assess tests, and analyze LLM bug-fixing performance in a flexible, remote contract role.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026