Contract JavaScript/TypeScript developer to build and evaluate AI models and datasets (20+ hrs/week, remote worldwide). Combine full-stack development with model benchmarking, response ranking, RLHF/SFT support, and written evaluation rationales.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Intermediate
Experience
Jul 17, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow freelance careers teaching AI by centralizing projects, tracking experience, and building a portable portfolio of model-evaluation and annotation work.
OpenTrain AI is the hiring and contracting organization for this role. Creating an OpenTrain account is free and is how you will apply and manage your work.
Why AI training matters
AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. Developers and evaluators create the examples and judgments modern models learn from—everything from response ranking to supervised fine-tuning and RLHF.
This field is highly flexible and remote-friendly, making it a great fit for freelancers who want part-time, impactful technical work that directly shapes model behavior.
The role
We are hiring a JavaScript/TypeScript Full-Stack Developer for freelance AI work focused on model training, evaluation, and response refinement. This is an intermediate-level contractor role expected to run at 20+ hours per week and is open to contributors worldwide who speak strong English.
Work blends hands-on development with model benchmarking, response ranking, dataset creation for SFT, and participation in RLHF workflows. You will write evaluation rationales and collaborate with cross-functional teams to improve models and product quality.
Language: English (strong spoken and written required)
What you'll do
You will write and maintain code that supports training and evaluating AI models, perform structured evaluations, produce clear rationales for judgments, and help create high-quality datasets for supervised fine-tuning and RLHF.
Expect to collaborate on code and documentation reviews and to partner with other contributors to analyze benchmark results and refine reward models and evaluation criteria.
Design, develop, and maintain code used to train and optimize AI models
Benchmark model performance and analyze results for improvement
Evaluate and rank model responses and write clear rationales
Support dataset creation for SFT and assist RLHF workflows
Review code and documentation; provide constructive technical feedback
Requirements
Candidates must have strong JavaScript or TypeScript experience and be comfortable building modular, scalable web apps. You should be able to write readable, reusable, and maintainable code and participate in code review processes.
A degree in Engineering, Computer Science, or equivalent experience is expected. Familiarity with testing, Docker, and software QA practices is a plus.
Strong JavaScript or TypeScript development experience
Experience with ES6 and at least one: Node.js, React, Nest.js, Angular, or Vue
Ability to write readable, reusable, maintainable code and perform code reviews
Comfort with testing, test planning, and Docker (preferred)
Strong spoken and written English communication skills
Bachelor’s or master’s degree in Engineering/Computer Science or equivalent experience
Helpful background
Prior exposure to AI model evaluations, response ranking, supervised fine-tuning (SFT), or reinforcement learning with human feedback (RLHF) will help you move faster in this role. Experience writing technical rationales for evaluation decisions is also valuable.
This role centers on TEXT data and label types RLHF, EVALUATION_RATING, and FINE_TUNING, so familiarity with those workflows is beneficial.
Experience with model evaluations, SFT, RLHF, or response ranking
Experience producing clear technical rationales for review decisions
Comfort working with text-centric datasets and evaluation labels
How it works & how to apply
OpenTrain AI hires and manages contracts for this role. Compensation details are not specified in this listing and will be provided during the application process.
To apply, create a free OpenTrain account, complete your profile with relevant JavaScript/TypeScript and AI evaluation experience, and submit your application through the OpenTrain platform.
Employment type: Contractor, part-time
Time commitment: 20+ hours per week
Data type: TEXT; label types: RLHF, EVALUATION_RATING, FINE_TUNING
Join OpenTrain as an AI Model Evaluation Developer to write and maintain code, run model benchmarks, rank responses, and build datasets for fine-tuning and RLHF. This remote, part-time contractor role requires strong Python and JavaScript/TypeScript skills and 20+ hours/week.
Audit AI-generated JavaScript/Angular submissions by building and running projects, checking functionality, security, performance, and accessibility, and correcting mis-ratings with concise feedback. Remote contractor, 20+ hrs/week at $23/hr; requires 7+ years professional Angular experience.
Contractor role reviewing AI-generated React and Next.js code, producing reference solutions and rating model responses; fully remote, 20+ hrs/week at $90/hr. Help improve AI training content for advanced frontend engineering topics.