Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Intermediate
Experience
Jul 17, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help freelancers discover projects, build a unified portfolio of AI training work, and grow durable freelance careers in a fast-moving space.
We hire and contract directly for specialized AI training roles. As an OpenTrain contributor you’ll work on real model-training tasks that shape how modern AI systems behave.
About AI training work
AI training (also called data labeling or human feedback work) is the human side of building AI: people create, evaluate, and refine the examples AI models learn from. This role focuses on evaluating and improving generative models and reward models using both automated and human-in-the-loop processes.
Typical benefits include fully remote, flexible schedules and hands-on influence over model behavior—ideal for developers who want to apply software skills directly to model quality and evaluation.
Role overview
You will be an AI Model Evaluation Developer responsible for writing and maintaining code used in evaluation and training workflows, running and analyzing model benchmark runs, ranking model outputs, and helping refine reward models with RLHF processes.
This is a contractor, part-time position requiring 20+ hours per week, remote work in English, and collaboration with cross-functional contributors on evaluation strategy and dataset creation.
What you'll do
Your day-to-day combines software development with model evaluation and dataset work. You’ll build and improve tools, run experiments, and write clear rationales that explain evaluation decisions.
Design, develop, and maintain efficient Python and JavaScript/TypeScript code used to train, evaluate, and optimize models.
Run evaluations and benchmark model performance; analyze results and recommend improvements.
Rank AI model responses to queries using predefined criteria and document scoring rationales.
Create and maintain high-quality datasets for supervised fine-tuning and evaluation.
Collaborate on RLHF workflows and reward-model refinement with human feedback.
Review code and documentation; provide constructive feedback to improve quality and stability.
Requirements
We require demonstrable software engineering skills, experience with evaluation tasks, and strong English communication for written rationales and collaboration.
Strong Python plus JavaScript or TypeScript development experience.
Proven ability to write readable, reusable, and maintainable code for web apps or scalable architectures.
Experience with code reviews and maintaining code quality standards.
Familiarity with web app development, modular design, and scalable architectures.
Understanding of testing, security, and stability practices.
Working knowledge of Docker.
Comfort evaluating model responses, ranking outputs, and explaining rationale clearly.
Ability to work 20+ hours per week as a contractor; English required.
Helpful background & nice-to-haves
These qualifications are valuable but not required. They help you move faster in our evaluation and engineering workflows.
Bachelor’s or Master’s degree in Engineering, Computer Science, or equivalent experience.
Experience with Node.js or Nest.js backend frameworks.
Experience with frontend frameworks such as React, Angular, or Vue.js.
Exposure to QA, test planning, or prompting for LLMs.
How the contracting process works
OpenTrain hires contractors directly. You will be onboarded to our evaluation projects, provided with task instructions and tooling, and asked to deliver code, evaluations, and documented rationales.
We hire globally; this role is remote and requires English for collaboration. Apply through OpenTrain to be considered and to build a profile that tracks your AI training work.
Employment: Contractor, part-time.
Time expectation: 20+ hours per week.
Work language: English. Worldwide applicants welcome.
Contract JavaScript/TypeScript developer to build and evaluate AI models and datasets (20+ hrs/week, remote worldwide). Combine full-stack development with model benchmarking, response ranking, RLHF/SFT support, and written evaluation rationales.
Join OpenTrain AI to evaluate LLM outputs in finance, design rubrics, and help shape model training and benchmarks. Part-time contractor role (<20 hrs/week), remote worldwide, paying $100/hr for finance professionals with 2+ years' experience.
Join OpenTrain to evaluate and improve LLM outputs for brand strategy and growth marketing, writing solutions and refining rubrics; remote (US), $60–$80/hr, default 35 hrs/week with a 20+ hrs/week expectation. Work as a contractor on RLHF and evaluation-rating tasks.