AI Evaluation Analyst — Rubric Design & Agent Assessment
Join OpenTrain to build evaluation tasks, prompts, and clear grading rubrics that measure AI agents on practical workflows; part-time remote work at $20–$35/hr for experienced writers and rubric designers. Work 20+ hours/week producing concise, structured reports and evolving benchmarks.
Generative AI & RLHF
100% Remote Hourly · $20–$35/hr
$20–$35/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jun 28, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow durable freelance careers teaching AI by consolidating opportunities, supporting portfolio building, and providing clear paths to more advanced evaluation and annotation work.
As the hiring organization for this role, OpenTrain offers flexible, remote projects at the cutting edge of how AI systems are trained and evaluated.
About AI training work
AI training (also called data labeling or human feedback) is the human side of building AI systems. Contributors create examples, rate outputs, and write evaluation materials that directly shape model behavior.
This work is often remote, flexible, and accessible to people with strong language skills and attention to detail. It’s an opportunity to influence how AI performs in real-world workflows while working part time or as a contractor.
The role
You will design self-contained evaluation tasks, prompts, supporting files, and grading rubrics to benchmark AI agents on administrative and workflow scenarios. The position focuses on precise written descriptions, unambiguous success/failure criteria, and iterative refinement of evaluation methods as project needs change.
Role type: Contractor, part-time
Time commitment: 20+ hours/week
Primary data type: Text (evaluation of text generation and ratings)
What you'll do
Your day-to-day will center on observing AI agent behavior, documenting findings in clear English, and converting those observations into repeatable evaluation tasks and rubrics. You will collaborate with other evaluators and refine frameworks from feedback.
Design evaluation tasks that test AI performance on administrative and workflow scenarios
Create written grading rubrics with unambiguous success and failure criteria
Document AI agent behavior in concise, precise English
Refine tasks and rubrics based on feedback and collaboration
Adapt evaluation frameworks across different domains as requirements change
Requirements
Candidates must already meet the following experience and skill expectations. We will not invent or assume qualifications beyond those listed here.
3+ years in roles requiring written precision and structured thinking
Native or fluent English writing with the ability to produce succinct, specific observations
Experience creating rubric-based evaluation, grading, or structured scoring frameworks
Strong attention to detail and ability to spot subtle inconsistencies
Comfort with computers, SaaS tools, web browsers, file management, and document editing
Self-direction and ability to work through ambiguous or loosely defined projects
Helpful background
The following experiences are not required but will help you succeed and move into more advanced evaluation work.
Prior experience evaluating AI outputs (text generation or rating tasks)
Experience refining evaluation rubrics or scoring methodologies
Ability to work across multiple domains and adjust quickly to new workflow challenges
Logistics, pay, and how we work
This is a remote contractor role that accepts applicants worldwide. You must be comfortable working independently and communicating findings in clear English.
Compensation is hourly: $20–$35 USD per hour (projects typically start at the lower end, with experienced evaluators eligible for higher rates based on skill and performance). Work is paid per hour and scheduled flexibly to meet the 20+ hours/week expectation.
Languages required: English (fluent/native)
Employment types: Contractor, Part-time
Pay type: Hourly (USD) — $20 minimum, up to $35 per hour
Data and label types: Text evaluation, evaluation ratings, text generation assessment
Who should apply and next steps
Apply if you enjoy tight, clear writing, designing measurable rubrics, and translating behavioral observations into repeatable evaluation materials. Ideal candidates are disciplined, detail-oriented, and experienced producing structured scoring systems.
To apply, submit work samples that demonstrate rubric creation, concise evaluation reports, or examples of scoring/grading frameworks. OpenTrain will review applications and provide the next steps for onboarding and initial tasks.
Design structured evaluation scenarios and gold-standard behaviors for LLM-based agents in a remote, part-time contractor role (20+ hrs/week). Pay $18–$24/hr; requires QA-style thinking, basic Python/JavaScript, and strong written English.
Create multi-turn conversations, rubrics, and evaluation assets for frontier LLMs while working remotely as a contractor 20+ hours/week. Rapid onboarding and clear specs; paid on a per-task/hour basis at $20–$30/hr.
Evaluate JSON-formatted AI model outputs against written task instructions and scoring rubrics in a short-term, US-only remote project; pay ranges $50–$175/hr with a default 40-hour weekly commitment and 20+ hours/week availability required.