OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect experienced contributors with meaningful, remote work that helps shape how state-of-the-art AI systems behave.
Creating an OpenTrain account is free. We help people discover projects, consolidate opportunities across platforms, and build a unified AI training portfolio they control.
About AI training and RLHF work
AI training (also called data labeling, annotation, or human feedback work) is the human side of building AI. Modern generative models learn from examples and evaluations prepared by people; RLHF (reinforcement learning from human feedback) relies on clear, consistent human judgments to guide model behavior.
As an AI Rater Guidelines Writer you’ll translate technical program needs into rater-ready instructions that enable reproducible, high-quality evaluations across domains.
Role overview
We are hiring an AI Rater Guidelines Writer to design rating guidelines and rubrics that eliminate ambiguity, resolve contradictions, and cover edge cases so human raters can evaluate model outputs reliably.
You will work closely with research and program teams to convert complex specifications from finance, retail, insurance, legal, and sports into precise, domain-appropriate guidance for raters working on RLHF and GenAI evaluation tasks.
What you'll do
Guide research and program teams to clarify ambiguous program requirements into structured, rater-ready instructions.
Design non-contradictory rating guidelines and rubrics that can be applied consistently by raters, including for edge cases.
Evaluate and revise draft guidelines for ambiguity, contradiction, and coverage gaps until escalation is rarely needed.
Collaborate with subject-matter experts and program leads to ensure accuracy and consistency across guideline sets.
Requirements
3+ years professional experience in linguistics, instructional design, technical writing, or a closely related field with a focus on GenAI/RLHF.
Direct experience writing or refining guidelines and rubrics for human raters in a GenAI or RLHF context.
Demonstrated ability to work across multiple subject-matter domains and translate nuance into clear instructions.
Proven track record of resolving ambiguity and contradiction in written specifications, with concrete before/after examples.
Ability to commit to a consistent schedule of at least 35 hours per week during weekdays.
Fluent written English and excellent written communication skills.
Work authorization in the United States (this role is US-only).
Helpful background
The following are not required but will make you more competitive and effective in this role.
Experience in instructional design or curriculum development for adult learners.
Familiarity with multiple domains beyond a single specialization, especially finance, legal, insurance, retail, or sports.
Previous work on AI training data quality, model evaluation, or related program design projects.
Compensation, schedule, and how it works
This is a contractor, part-time engagement with a required weekday schedule. Minimum commitment: at least 35 hours/week.
Pay is hourly at $45–$65 USD per hour, paid per the project’s contracting terms.
Labeling work will focus on text and RLHF-style rating tasks. You’ll collaborate with program teams and subject-matter experts and deliver polished guideline documents and rubrics.
Employment type: Contractor, Part-time.
Language: English required. Location: United States only.
Primary data type: Text. Label type: RLHF (generation evaluation and rating).
Next steps
If this matches your skills and availability, create or use your OpenTrain profile to apply. Include examples showing before/after guideline edits or rubrics you authored and a summary of domain experience.
We review applications for clear relevant experience, demonstrated writing samples, and the ability to maintain a consistent weekday schedule.
Contract role evaluating and improving Latin–English AI-generated text for model training and RLHF; BA/BS required. Fixed-price project of $8,000, remote work open to candidates in the US, Canada, Europe, and Latin America.
Create multi-turn conversations, rubrics, and evaluation assets for frontier LLMs while working remotely as a contractor 20+ hours/week. Rapid onboarding and clear specs; paid on a per-task/hour basis at $20–$30/hr.
Work remotely as a Hebrew–English bilingual evaluator, reviewing and improving AI-generated text for accuracy, clarity, and reasoning. Part-time contractor role at $32/hr (under 20 hrs/week) for experienced translators, editors, or linguistic QA specialists.