Evaluate ChatGPT and Claude responses, create challenging prompts, and explain model strengths and weaknesses in clear American English. Work remotely for 20+ hours per week at $20-$30 per hour.
Generative AI & RLHF
100% Remote Hourly · $20–$30/hr
$20–$30/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 15, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting contributors for practical projects that help improve how modern artificial intelligence systems understand and generate language.
Build experience in a fast-growing AI training industry
Work remotely from anywhere in the world
Create a free OpenTrain account and apply in minutes
About AI Model Evaluation
AI models learn from human-created examples, comparisons, ratings, and feedback. In evaluation work, contributors assess model responses against clear guidelines and provide careful judgments that help make AI systems more accurate, useful, and consistent.
Support the development of generative AI through human feedback
Use critical thinking and communication skills in hands-on model review
Work part time with a schedule of 20 or more hours per week
The Role
OpenTrain is seeking an English LLM Evaluation Generalist to create prompts, compare large language model responses, and provide structured verbal feedback. You will work across a wide range of topics, evaluate outputs from ChatGPT and Claude in real time, and record your screen and microphone during each session.
This entry-level contractor role is suited to someone with strong American English communication, practical experience using generative AI tools, and the judgment to explain model strengths, weaknesses, and meaningful differences.
Employment type: Part-time contractor
Experience level: Entry level
Time requirement: 20+ hours per week
Pay: $20-$30 per hour
Location: Worldwide and fully remote
Language: Fluent American English
What You'll Do
You will evaluate language model behavior through prompt creation, side-by-side comparison, and detailed feedback. Each session requires focused work in a distraction-free setting, consistent use of evaluation guidance, and clear communication of your reasoning.
Create prompts that challenge large language models across varied topics
Compare responses from ChatGPT and Claude
Identify strengths, weaknesses, and meaningful differences between responses
Record your screen and microphone while completing evaluation sessions
Deliver detailed verbal feedback
Interpret guidance documentation and apply consistent evaluation standards
Maintain clarity, professionalism, and accuracy throughout each session
Requirements
Applicants should be fluent in spoken and written American English and comfortable explaining preferences and judgments clearly. You should also be able to work reliably with the required recording setup and apply detailed instructions consistently.
High-level fluency in spoken and written American English
Experience using ChatGPT, Claude, Gemini, or similar large language models
Familiarity with AI evaluation, content evaluation, data annotation, or quality assurance
Strong critical thinking and analytical reasoning
Excellent attention to detail
Reliable computer and stable internet connection
Ability to record high-quality audio and video without issues
Comfort comparing ChatGPT and Claude responses
Ability to explain model strengths, weaknesses, and preferences clearly
Helpful Background
Research experience or other work involving structured comparison and written or verbal analysis can help you succeed in this role. The work centers on practical judgment, careful reasoning, and consistent communication.
Research experience
Structured comparison or analysis experience
Clear written and verbal communication
How to Apply Through OpenTrain
Create a free OpenTrain account to build your AI training profile and apply for this opportunity. OpenTrain brings contributors into a growing field where human reviewers help shape the behavior of cutting-edge AI systems.
Review the role requirements
Create or update your OpenTrain profile
Apply through OpenTrain in minutes
Prepare a reliable computer, internet connection, and recording setup
Help improve large language models by evaluating AI-generated responses, researching claims, analyzing data, and writing clear feedback. This entry-level contractor role offers remote, part-time work of 20+ hours per week.
Help improve large language models by creating challenging biology problems, writing rigorous solutions, and evaluating model reasoning from undergraduate through PhD level. This remote expert contract requires 20+ hours weekly.
Use deep sports knowledge to write challenging prompts, evaluate large language model responses, and identify factual or reasoning issues. This remote contractor role offers 20+ hours per week for experts with a master's degree and three years of relevant experience.