Skip to content
OpenTrain AIFor AI Companies

AI Model Evaluation Developer

Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Intermediate

Experience

Jul 17, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help freelancers discover projects, build a unified portfolio of AI training work, and grow durable freelance careers in a fast-moving space.

We hire and contract directly for specialized AI training roles. As an OpenTrain contributor you’ll work on real model-training tasks that shape how modern AI systems behave.

About AI training work

AI training (also called data labeling or human feedback work) is the human side of building AI: people create, evaluate, and refine the examples AI models learn from. This role focuses on evaluating and improving generative models and reward models using both automated and human-in-the-loop processes.

Typical benefits include fully remote, flexible schedules and hands-on influence over model behavior—ideal for developers who want to apply software skills directly to model quality and evaluation.

Role overview

You will be an AI Model Evaluation Developer responsible for writing and maintaining code used in evaluation and training workflows, running and analyzing model benchmark runs, ranking model outputs, and helping refine reward models with RLHF processes.

This is a contractor, part-time position requiring 20+ hours per week, remote work in English, and collaboration with cross-functional contributors on evaluation strategy and dataset creation.

What you'll do

Your day-to-day combines software development with model evaluation and dataset work. You’ll build and improve tools, run experiments, and write clear rationales that explain evaluation decisions.

  • Design, develop, and maintain efficient Python and JavaScript/TypeScript code used to train, evaluate, and optimize models.
  • Run evaluations and benchmark model performance; analyze results and recommend improvements.
  • Rank AI model responses to queries using predefined criteria and document scoring rationales.
  • Create and maintain high-quality datasets for supervised fine-tuning and evaluation.
  • Collaborate on RLHF workflows and reward-model refinement with human feedback.
  • Review code and documentation; provide constructive feedback to improve quality and stability.

Requirements

We require demonstrable software engineering skills, experience with evaluation tasks, and strong English communication for written rationales and collaboration.

  • Strong Python plus JavaScript or TypeScript development experience.
  • Proven ability to write readable, reusable, and maintainable code for web apps or scalable architectures.
  • Experience with code reviews and maintaining code quality standards.
  • Familiarity with web app development, modular design, and scalable architectures.
  • Understanding of testing, security, and stability practices.
  • Working knowledge of Docker.
  • Comfort evaluating model responses, ranking outputs, and explaining rationale clearly.
  • Ability to work 20+ hours per week as a contractor; English required.

Helpful background & nice-to-haves

These qualifications are valuable but not required. They help you move faster in our evaluation and engineering workflows.

  • Bachelor’s or Master’s degree in Engineering, Computer Science, or equivalent experience.
  • Experience with Node.js or Nest.js backend frameworks.
  • Experience with frontend frameworks such as React, Angular, or Vue.js.
  • Exposure to QA, test planning, or prompting for LLMs.

How the contracting process works

OpenTrain hires contractors directly. You will be onboarded to our evaluation projects, provided with task instructions and tooling, and asked to deliver code, evaluations, and documented rationales.

We hire globally; this role is remote and requires English for collaboration. Apply through OpenTrain to be considered and to build a profile that tracks your AI training work.

  • Employment: Contractor, part-time.
  • Time expectation: 20+ hours per week.
  • Work language: English. Worldwide applicants welcome.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

JavaScript / TypeScript AI Evaluation Developer

Contract JavaScript/TypeScript developer to build and evaluate AI models and datasets (20+ hrs/week, remote worldwide). Combine full-stack development with model benchmarking, response ranking, RLHF/SFT support, and written evaluation rationales.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Intermediate level

Posted Jul 17, 2026

Finance Model Evaluation Expert

Join OpenTrain AI to evaluate LLM outputs in finance, design rubrics, and help shape model training and benchmarks. Part-time contractor role (<20 hrs/week), remote worldwide, paying $100/hr for finance professionals with 2+ years' experience.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $100/hr

Posted Jul 16, 2026

Marketing AI Model Evaluation Expert

Join OpenTrain to evaluate and improve LLM outputs for brand strategy and growth marketing, writing solutions and refining rubrics; remote (US), $60–$80/hr, default 35 hrs/week with a 20+ hrs/week expectation. Work as a contractor on RLHF and evaluation-rating tasks.

Generative AI & RLHF
Text
Remote · United States
English
Part-time · Flexible
Expert level
Hourly · $60–$80/hr

Posted Jul 13, 2026