Use strong Python skills to create coding data, evaluate language-model responses, and improve AI training workflows. This fully remote, one-month contractor assignment offers 20, 30, or 40 hours per week.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this role, giving contributors a way to build experience in a fast-growing field where human expertise helps shape advanced AI systems.
Fully remote work available worldwide
Free OpenTrain account and profile
Build a portfolio around AI training and data-labeling experience
About AI Model Evaluation Work
AI training is the human side of building artificial intelligence. Contributors create examples, review model outputs, write evaluations, and provide feedback that helps language models become more accurate, useful, and aligned.
This role focuses on coding and evaluation rather than building or fine-tuning the models themselves. Your Python expertise and technical judgment will help produce reliable training data and assess the quality of model-generated solutions.
Create supervised fine-tuning examples
Evaluate and rank language-model responses
Support reinforcement learning from human feedback workflows
Help improve the quality of AI-generated code and technical explanations
The Role
OpenTrain is seeking a Python AI Model Evaluation Developer to support large language model improvement through hands-on coding, data generation, and evaluation. You will write Python solutions to code-based questions, create high-quality training examples, compare responses from different models, and provide detailed feedback.
This is a fully remote contractor assignment lasting one month. The role is categorized as entry level, while the stated technical requirements call for at least three years of strong Python programming experience.
Engagement: Contractor and part time
Duration: One month
Workload options: 20, 30, or 40 hours per week
Minimum commitment: 20 hours per week and at least 4 hours per day
Required overlap: 4 hours with Pacific Time
Location: Worldwide and fully remote
Working language: Fluent written and conversational English
Compensation: Not disclosed
What You’ll Do
You will combine software-development discipline with careful model evaluation. The work includes creating and reviewing technical content, analyzing performance, and delivering feedback that researchers and annotators can use to strengthen AI training processes.
Design, develop, and maintain efficient, high-quality Python code for AI training and evaluation workflows.
Conduct evaluations to benchmark model performance and analyze results for continuous improvement.
Evaluate and rank AI model responses using quality, relevance, accuracy, and alignment criteria.
Create task-specific supervised fine-tuning datasets, model responses, code solutions, and evaluation rationales.
Create and refine responses to improve clarity, relevance, and technical accuracy.
Contribute to reinforcement learning from human feedback activities and reward-model refinement.
Design evaluation strategies, review code and documentation, and provide constructive technical feedback.
Collaborate with researchers and annotators while exploring tools and methods that strengthen AI training processes.
Requirements
Applicants should bring strong Python programming ability, software-development fundamentals, and the communication skills needed to explain technical judgments clearly. Experience working with model responses and AI training data-generation workflows is also required.
At least three years of strong Python programming experience
Familiarity with Python frameworks and libraries
Knowledge of software-development quality, formatting, architecture, and best practices
Experience with unit, integration, and property-based testing in Python
Understanding of multithreading and asynchronous programming
Ability to refactor code safely without introducing regressions
Ability to diagnose memory or concurrency issues
Ability to write clear evaluation rationales and technically accurate responses
Working knowledge of supervised fine-tuning and reinforcement learning from human feedback data-generation and evaluation workflows
Fluent conversational and written English communication
Who Should Apply
This opportunity is suited to Python developers who enjoy solving technical problems, reviewing code, and making precise quality judgments. It may also appeal to software engineers interested in applying their skills to the rapidly growing field of AI training.
Attention to detail matters because your code, rankings, rationales, and feedback will be used to assess and improve language-model behavior. The assignment requires a consistent weekly commitment and Pacific Time overlap.
Python developers with strong testing and debugging experience
Software professionals comfortable with concurrency and asynchronous programming
Technical communicators who can explain evaluation decisions clearly
Contributors interested in coding datasets, model evaluation, and RLHF work
How It Works
Apply through OpenTrain to be considered for this contractor assignment. If selected, you will complete remote AI training and evaluation work within the available 20, 30, or 40 hour weekly commitment.
OpenTrain helps contributors discover and grow careers in AI training and data labeling. Your profile can showcase relevant experience and support a longer-term portfolio as you take on additional opportunities in the field.
Create a free OpenTrain account
Build a profile highlighting Python and evaluation experience
Apply in minutes through OpenTrain
Complete the assignment remotely according to the required schedule
Evaluate next-generation AI coding agents by reviewing their workflows, debugging Python outputs, and validating technical results. This remote contractual role requires expert Python engineering experience and practical experience building or using LLM-powered agents.
Use your Python and AI engineering experience to evaluate coding-agent trajectories, tool calls, code changes, and technical outcomes. Work worldwide on a flexible 20+ hour-per-week contract through OpenTrain.
Evaluate Python coding agents by reviewing trajectories, tool use, code changes, execution results, and final outputs. This remote contractor role offers 20+ hours per week for experienced software engineers with AI or LLM expertise.