Help improve large language models by reviewing English response data, queries, context, and auto-generated labels. This worldwide, part-time contractor role pays $5 per hour and uses a structured three-person quality process.
Generative AI & RLHF
100% Remote Hourly · $5/hr
$5/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Oct 31, 2024
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover opportunities to teach and improve modern AI systems.
Worldwide opportunity
Part-time contractor engagement
Apply and build your AI training career through OpenTrain
About AI Training Work
AI training is the human work behind today’s artificial intelligence systems. Contributors review examples, classify information, evaluate model responses, and provide quality judgments that help AI models become more useful and reliable.
This role focuses on text-based English documents and large language model responses. It is remote work that can fit around other commitments while giving you hands-on experience with cutting-edge AI development.
Work with language-model response data
Support the development of more reliable AI systems
Use human judgment to verify automated labels
The Role
OpenTrain AI is seeking an expert-level LLM Data Labelling and Annotation contractor to review approximately 500 to 1,000 rows of data. Each row includes context, a query, and auto-generated labels that require human verification.
The work involves English-language content and may include technical computer science or financial material. Reviewers should use best-effort judgment on specialized items, with expert review supporting technical and financial content.
Role: LLM Data Labelling and Annotation
Data type: Text-based English documents
Engagement: Part-time contractor
Time requirement: 20 or more hours per week
Pay: $5 per hour
What You'll Do
You will assess LLM response data and verify whether the associated auto-generated labels are accurate and appropriate. The project uses a structured two-step human review process designed to create consistent, carefully checked training data.
Review context, queries, and LLM-generated responses
Verify auto-generated labels across approximately 500 to 1,000 rows
Perform classification and entity or named-entity classification
Evaluate and rate response quality
Apply English-language understanding and basic general knowledge
Make best-effort judgments on technical and financial content when needed
Review Process and Quality Assurance
The first review step assigns the exact same row to two agents so their judgments can be compared. A third agent then performs quality assurance. At least one of the three agents must be an expert capable of assessing technical and financial content.
Two agents review the same row during the first step
A third agent completes quality assurance
Technical and financial items receive expert-supported assessment
Label types include classification, entity NER classification, and evaluation rating
Requirements
Familiarity with the English language is required, along with basic general knowledge. Experience working with technical computer science data or financial documents is a plus. The project is listed at an expert experience level, particularly for reviewers assessing specialized content.
Familiarity with English is required
Basic general knowledge
Technical computer science data experience is a plus
Financial document experience is a plus
Ability to assess specialized content when assigned
Availability for 20 or more hours per week
Work Details
This is a worldwide, part-time contractor opportunity with hourly pay of $5 USD. Labeling work is completed using AWS SageMaker, and the project is centered on text data and human evaluation of LLM outputs.
Create analytical questions, scenarios, explanations, and text-based training data that help improve large language models. This worldwide, remote contractor role offers flexible part-time work of 20+ hours per week.
Short contract for a bilingual (German/English) reviewer to label and create analytical text scenarios that improve LLM reasoning. One-week, 40 hrs/week remote role with required daily overlap to UTC-8; immediate start.
Short-term remote contract for a bilingual French/English analyst to create, label, and evaluate analytical Q&A that improve large language models; part-time (20+ hrs/week) with requirements to overlap US hours and potential extension. Entry-level friendly for strong analytical writers.