Use advanced Python skills to evaluate AI models, create training data, rank responses, and improve coding-focused systems. This worldwide contract role offers flexible part-time work of 20+ hours per week.
Coding & Software
100% Remote
Worldwide
Eligibility
Intermediate
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the organization hiring and contracting for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free, and this role is open to qualified contractors worldwide.
About AI Model Evaluation Work
AI training is the human work behind modern artificial intelligence. Contributors write prompts, create and review model responses, evaluate technical accuracy, and provide structured feedback that helps AI systems become more useful and reliable.
In this project, your Python expertise will support coding-focused evaluation, supervised fine-tuning data, and RLHF-style feedback workflows.
The Role
OpenTrain AI is seeking a Senior Python Developer for AI model evaluation and data generation work. You will combine software development experience with careful judgment to assess multiple large language model outputs and create high-quality examples for training and evaluation.
This is an intermediate-level, part-time contractor role requiring 20+ hours per week. The work is worldwide and conducted in English.
Employment type: Contractor
Schedule: Part time, 20+ hours per week
Location: Worldwide
Working language: English
What You'll Do
You will write, review, and explain code-based solutions while helping develop evaluation methods for AI models. Clear reasoning and consistent application of evaluation criteria are central to the work.
Write efficient Python code for AI training and evaluation workflows
Build responses and solutions for prompt-based coding tasks
Evaluate and rank model outputs against predefined criteria
Produce rationales and peer reviews that support model improvement
Contribute to supervised fine-tuning datasets and RLHF-style feedback workflows
Review code and documentation for quality, correctness, and regressions
Work with researchers, annotators, and cross-functional teams on model performance
Required Skills and Experience
You should have strong Python development experience and be comfortable assessing both code quality and AI-generated responses. The role requires clear written and conversational English communication.
3+ years of strong Python programming experience
Experience with unit, integration, and property-based testing
Knowledge of multithreading and asynchronous programming in Python
Ability to refactor code safely and debug memory or concurrency issues
Familiarity with code quality, formatting, and software development best practices
Experience writing code for AI training or evaluation workflows
Ability to compare model outputs and explain evaluation decisions
Familiarity with SFT, RLHF, and benchmark-style model assessment
Fluent written and conversational English
Why This Work Matters
Every major AI system depends on people who prepare examples, review outputs, and identify where models need to improve. By combining your Python skills with thoughtful evaluation, you can contribute directly to how advanced AI systems behave while working remotely in a growing technical field.
Apply software engineering expertise to cutting-edge AI development
Help improve model alignment, clarity, and technical accuracy
Work remotely with a flexible part-time contractor schedule
Build experience in AI training, evaluation, and data generation
Build and evaluate AI systems through Python development, supervised fine-tuning, RLHF, response ranking, and public-data analysis. This fully remote contractor role offers 20, 30, or 40 hours per week for a one-month engagement.
Build scalable Python APIs and help test AI coding tools in focused four-day bursts. This remote contractor role offers $50–$100 per hour, 20+ hours weekly, and the chance to shape developer-focused AI systems.
Help evaluate how large language models work with real Python code. Build repository environments, triage issues, assess tests, and analyze LLM bug-fixing performance in a flexible, remote contract role.