Build and evaluate AI systems through Python development, supervised fine-tuning, RLHF, response ranking, and public-data analysis. This fully remote contractor role offers 20, 30, or 40 hours per week for a one-month engagement.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build professional profiles, and grow in a rapidly expanding field.
About AI Training and Model Evaluation
AI training is the human side of building modern artificial intelligence. Specialists prepare datasets, evaluate model outputs, write and rank responses, and provide feedback that helps AI systems become more capable, accurate, and useful.
Work on cutting-edge AI development with fully remote flexibility.
Help shape how advanced models perform through evaluation, fine-tuning, and human feedback.
Build experience in a fast-growing technical field alongside other contract and part-time opportunities.
The Role
OpenTrain AI is seeking a Python AI Model Training & Evaluation Engineer for a fully remote contractor engagement. You will work with US companies on advanced commercial and research AI solutions, combining Python engineering, model evaluation, dataset development, supervised fine-tuning, reinforcement learning with human feedback, and public-data analysis.
This is an entry-level role with a minimum commitment of 20 hours per week. You may choose 20, 30, or 40 hours per week, provided that at least four hours per day overlap with Pacific Standard Time.
Engagement type: Contractor and part-time
Duration: One month
Start: Next week
Location: Fully remote, worldwide
Schedule: 20, 30, or 40 hours per week
Time-zone overlap: At least 4 hours per day with PST
What You'll Do
You will contribute to the full AI model training and evaluation workflow, from writing Python code and preparing task-specific datasets to reviewing model responses and communicating analytical findings. The work requires careful judgment, clear documentation, and the ability to explain technical conclusions to researchers and stakeholders.
Write Python code to train, optimize, and evaluate AI models.
Benchmark model performance through evaluations, or Evals.
Evaluate and rank AI model responses to user queries across diverse domains.
Provide detailed rationales for response rankings and evaluation decisions.
Lead supervised fine-tuning efforts by creating and maintaining high-quality, task-specific datasets.
Collaborate on reinforcement learning with human feedback to refine reward models.
Analyze public datasets from sources such as Kaggle, the UN, and US government datasets to answer business queries.
Document findings and communicate complex conclusions clearly.
Conduct peer reviews of code and documentation and provide constructive feedback.
Required Qualifications
This role is suited to someone with strong Python and analytical capabilities who can work carefully with AI model outputs and communicate findings in fluent English. A bachelor's or master's degree in Engineering, Computer Science, or equivalent experience is required.
Strong proficiency in Python, data analysis, and data science.
Excellent problem-solving and analytical skills.
Fluent conversational and written English.
Ability to communicate complex findings clearly to researchers and stakeholders.
Ability to evaluate and rank AI model responses with clear rationales.
Strong analytical skills for drawing conclusions from public datasets.
Availability for at least 20 hours per week with the required PST overlap.
Helpful Background
Experience in the following areas will help you contribute quickly, although the role is listed at the entry level:
Supervised fine-tuning, reinforcement learning with human feedback, or AI model evaluation and ranking.
Developing training datasets or evaluating model responses.
Using Jupyter notebooks.
Analyzing public datasets, including data from Kaggle, the UN, or US government sources.
Python programming and data analysis for AI model training and evaluation.
Why Join AI Training Work
AI training connects technical problem-solving with the development of state-of-the-art systems. Remote contributors can build practical experience in model behavior, data quality, evaluation, and human feedback while choosing work that fits their schedule.
Fully remote work from anywhere with a computer and internet connection.
A flexible part-time structure with options for 20, 30, or 40 hours per week.
Direct involvement in model training, evaluation, and dataset quality.
An opportunity to strengthen experience across Python, data science, and machine learning.
Use advanced Python skills to evaluate AI models, create training data, rank responses, and improve coding-focused systems. This worldwide contract role offers flexible part-time work of 20+ hours per week.
Join OpenTrain as a remote Machine Learning Engineer focused on benchmark-driven evaluation of real-world ML systems. This contractor role requires 3+ years of ML engineering experience, strong Python skills, and availability 20+ hrs/week with PST overlap.