Join OpenTrain to evaluate large language model outputs, create challenging prompts, and deliver recorded verbal feedback; remote, contract role (20+ hrs/week) paying $20–$30/hr. Entry-level friendly for strong American English speakers with LLM experience.
Generative AI & RLHF
100% Remote Hourly · $20–$30/hr
$20–$30/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 15, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help freelancers discover specialized AI training work, build a unified portfolio, and grow durable freelance careers in an expanding field.
This role is hired and contracted through OpenTrain AI — you will work directly with our team on model-evaluation projects and have your contributions tracked in your OpenTrain profile.
About AI training work
AI training (also called data labeling or human feedback work) is the human side of building modern AI: people create, evaluate, and refine examples that teach models how to behave. These projects are mostly remote, flexible, and accessible to contributors with clear communication and attention to detail.
As an evaluator you will directly shape how state-of-the-art language models perform by crafting prompts, comparing outputs, and providing structured feedback.
The role
You will serve as an English LLM Evaluation Generalist focused on prompt creation, side-by-side model comparison, and structured verbal feedback. Work is text-focused and will include evaluation rating and text-generation tasks.
This is a contract, part-time role expecting 20+ hours per week; pay is $20–$30 per hour. The position is entry level but requires fluency in American English and familiarity with large language models.
Work type: Contractor, Part-time
Time commitment: 20+ hours/week
Pay: USD $20–$30 per hour
Data type: Text; label types: EVALUATION_RATING, TEXT_GENERATION
What you'll do
You will create diverse prompts, run them through LLMs, compare outputs side-by-side, and record detailed spoken evaluations while capturing your screen and microphone.
Accuracy, clear reasoning, and consistent application of evaluation guidelines are essential — you will explain strengths, weaknesses, and meaningful differences between model responses.
Write prompts that test models across varied topics and styles.
Compare responses from major LLMs (for example, ChatGPT and Claude) and highlight differences in quality and behavior.
Record high-quality screen and audio during sessions and deliver structured verbal feedback in a distraction-free environment.
Apply provided evaluation guidelines consistently and document your judgments clearly.
Requirements
You must meet all core requirements below to perform the role reliably and professionally.
Technical readiness — a reliable computer, stable internet, and the ability to record high-quality audio and video — is non-negotiable because sessions are recorded.
Fluent spoken and written American English.
Experience using large language models such as ChatGPT, Claude, Gemini, or similar tools.
Familiarity with AI evaluation tasks, data annotation, QA, or content evaluation.
Strong critical thinking, analytical reasoning, and attention to detail.
Reliable computer, stable internet connection, and the ability to record screen and microphone with clear audio.
Comfort explaining model strengths, weaknesses, and preferences clearly in spoken feedback.
Helpful background
You don't need deep research credentials, but experience that demonstrates structured comparison and clear analysis will help you succeed.
Relevant experience examples include prior AI evaluation, QA work, research tasks, or roles requiring systematic written or verbal analysis.
Research experience or structured analysis roles are a plus.
Previous data-labeling or model-evaluation work is helpful but not required.
Join OpenTrain AI to design, solve, and evaluate challenging mathematics problems that probe LLM reasoning limits; this remote, part-time contractor role expects 20+ hours/week, advanced graduate-level math knowledge, Python ability, and English fluency.
Join OpenTrain as an Insurance LLM Evaluation SME to design and score underwriting, claims, and risk-assessment evaluation tasks for LLMs. Remote (U.S. only), $60–$80/hr, 35 hours/week, contractor/part-time.
Design structured evaluation scenarios and gold-standard behaviors for LLM-based agents in a remote, part-time contractor role (20+ hrs/week). Pay $18–$24/hr; requires QA-style thinking, basic Python/JavaScript, and strong written English.