OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect people with real, paid work where they teach and shape how modern AI systems behave.
OpenTrain AI is the hiring and contracting organization for this role. Contributors work remotely, build skills that travel across the industry, and take on flexible projects that influence state-of-the-art models.
Why AI training matters
AI training (data labeling / human feedback) is the human side of building intelligent systems. People prepare and review examples—writing prompts, rating outputs, structuring inputs, and testing integrations—that modern models learn from.
This work is often remote, flexible, and accessible: it’s a way to do cutting-edge technical work without a traditional software-engineering job title.
The role
You will design, test, and refine LLM prompts and evaluation criteria; build real-world automation pipelines that use LLMs; and review generated outputs for quality and reliability in content generation, communication, and personalized outreach flows.
Time commitment: 20+ hours per week (part-time contractor).
Pay: hourly, USD $15–$45 per hour (listed top rate up to $45/hr).
Level: Intermediate.
Data type: Text; label types include EVALUATION_RATING, FINE_TUNING, RLHF, COMPUTER_PROGRAMMING_CODING, FUNCTION_CALLING.
What you'll do
Design and iterate instruction prompts, chained prompts, and structured-output templates to improve model responses.
Integrate LLMs into automation workflows via REST APIs, SDKs, webhooks, or automation tools to power content and communication flows.
Define and apply rubric-based evaluation criteria to rate outputs for quality, safety, and task-specific success.
Perform hands-on text annotation, evaluation, and quality assurance to support fine-tuning and RLHF pipelines.
Structure and clean inputs, build templates, and prepare data for reliable model conditioning and function-calling.
Review model behavior for limitations and propose mitigation strategies to improve robustness and reliability.
Requirements
Candidates must meet the experience and application requirements below; preserve requested application details when you apply.
Experience evaluating LLM outputs for legal reasoning quality.
Hands-on experience with text annotation, evaluation, or rubric-based QA.
Experience integrating LLM APIs via REST APIs, SDKs, or automation tools.
Strong prompt engineering background: instruction design, chaining, and structured outputs.
Skilled in building AI-powered content automation for communication and personalized outreach.
Proficient at data structuring, templating, and input cleaning.
Knowledge of LLM limitations and mitigation methods.
Experience building or operating workflows using APIs, webhooks, or automation tools (example: Zapier).
CV must be in English and indicate your level of English proficiency; include an email address and phone number on your CV.
Who should apply
This role is a fit for technically minded prompt engineers, ML engineers or product-focused practitioners who have hands-on experience evaluating outputs, building prompt-driven automations, and integrating models via APIs. If you have worked on fine-tuning, RLHF, or evaluation pipelines and enjoy iterating on prompts and workflows, apply.
Intermediate experience level expected; formal CS degrees are not required if you can show relevant hands-on results.
This role suits people who want part-time, remote contractor work that directly shapes model behavior.
How to apply and important restrictions
Apply through your OpenTrain profile. Include a CV in English that states your English proficiency level and provides an email address and phone number. OpenTrain AI will review applications and follow up with assessment tasks and a short technical screening.
Note on location restrictions: acquisition is restricted in the following places. Applicants located in these countries, territories, or states are not eligible:
Also restricted: Switzerland; China; Taiwan; Kenya.
Restricted US states: Alaska, Arkansas, California, Connecticut, Delaware, Georgia, Hawaii, Illinois, Indiana, Kansas, Louisiana, Maine, Maryland, Massachusetts, Nebraska, Nevada, New Hampshire, New Jersey, New Mexico, Ohio, Oregon, Tennessee, Utah, Vermont, Washington, West Virginia.
Additionally restricted territories: Antarctica, Aruba, Åland Islands, Saint Barthélemy, Bonaire/Sint Eustatius and Saba, Bouvet Island, Cocos (Keeling) Islands, Democratic Republic of the Congo, Cook Islands, Christmas Island, Western Sahara, Falkland Islands (Malvinas), French Guiana, Guadeloupe,
Write realistic multi-turn conversations that teach assistants when and how to call tools, modeling calendars, email, maps, and other app workflows. Part-time contractor role requiring 20+ hrs/week and 3+ years technical or analytical experience; open to applicants in specified countries.
Lead large technical teams to deliver high-quality SFT and RLHF datasets for foundational language models. This part-time contractor role requires senior engineering leadership, hands-on quality judgment, and 20+ hours/week from candidates based in India, Brazil, Mexico, Argentina, Chile, Colombia,
Join OpenTrain to audit and improve LLM prompts and model answers on vehicle systems—verify calculations, rate against rubrics, and draft better solutions. Part-time contractor role, $40/hr, under 20 hours/week; requires 3+ years in automotive engineering and practical Python.