Contract role to design and evaluate computer-science problems and validate AI solutions using Python (NumPy/Pandas/SciPy). Remote, project-based work paying $15–$40/hr with a typical commitment of ~10–20 hours/week during active phases.
Generative AI & RLHF
100% Remote Hourly · $15–$40/hr
$15–$40/hr
Compensation
Worldwide
Eligibility
Expert
Experience
Apr 5, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect skilled contributors with projects that shape how state-of-the-art AI systems behave, offering flexible, remote work and a path to grow technical expertise in a fast-moving industry.
Why AI training matters
AI training (also called data labeling, annotation, or human feedback) is the human side of building AI. Contributors create, evaluate, and refine examples that models learn from — work that’s remote, flexible, accessible, and directly impacts real systems.
Work remotely and flexibly around other commitments.
Contribute to cutting-edge AI behavior by validating model reasoning and correctness.
The role
We’re hiring Computer Science experts to design rigorous problems and evaluate AI-generated solutions with strong technical rigor. You will validate multi-step reasoning and use Python (NumPy, Pandas, SciPy or equivalents) to check calculations, simulations, and correctness.
This is a contract, part-time, project-based role. Work is remote and open globally; screening notes include a location restriction referencing the USA (see Requirements).
Labeling types: evaluation/rating, question answering, text generation, prompt+response (SFT), programming/coding validation, and red-teaming.
Data type: text. Pre-labeled data: no. /unspecified.
What you’ll do day-to-day
Design real-world computer science problems, review AI solutions for correctness, and score multi-step reasoning using structured criteria. Use Python to reproduce, validate, or simulate results when needed, and improve explanations so AI reasoning aligns with industry standards.
You’ll also participate in expert discussions to maintain consistent evaluation standards and contribute to problem sets that reflect professional practice.
Validate calculations, algorithms, and simulations using Python (NumPy, Pandas, SciPy or equivalent).
Apply structured scoring to assess correctness, assumptions, constraints, and edge cases.
Write and refine prompts and model responses (SFT) and perform red-team style evaluations for robustness.
Requirements and screening
Candidates must demonstrate strong CS expertise, Python proficiency, and excellent technical writing in English. Preserve the listed screening details when applying.
Degree in Computer Science or closely related field, or equivalent demonstrated expertise.
2+ years of applied, research, or teaching experience in CS or adjacent technical fields.
Strong Python proficiency for numerical validation/simulation (NumPy, Pandas, SciPy or equivalents).
Experience with structured evaluation/scoring of multi-step reasoning problems.
Written English at C1+ level (clear technical writing and feedback).
Availability to contribute approximately 10–20 hours/week during active project phases; some projects may request 20+ hours/week.
Familiarity with other languages/tools (MATLAB, R, C/C++, SQL, domain-specific libraries) is a plus.
Professional certifications and international/applied project experience are advantageous.
Location details: listing notes 'Global — Any Location' and also references 'Restricted Location: USA' in screening; please confirm eligibility during application.
Who should apply
Apply if you are an experienced computer scientist, instructor, researcher, or practitioner comfortable judging multi-step reasoning and reproducing results with Python. This role fits contributors who want flexible, impactful work shaping how AI reasons about technical problems.
You enjoy precise technical review, clear feedback, and working with expert peers.
You want project-based, remote work and can commit to the stated weekly availability during active phases.
How it works & compensation
Assignments are project-based and typically run in active phases where focused weekly contribution is required. OpenTrain manages hiring and contracting for this role; contributors are engaged as contractors on part-time terms.
Compensation is hourly. The posting lists a pay range of $15–$40 per hour depending on task complexity and expertise.
Employment type: Contractor, Part-time.
Pay type: Hourly (PAY_PER_HOUR). Range: $15 to $40 / hour.
Typical time expectation: ~10–20 hours/week during active phases; some tasks may require up to or over 20 hours/week.
Design original civil/computational engineering problems and verify reproducible solutions using Python to create high-quality training data for generative models. Part-time contract (20+ hrs/week), up to $50/hr; requires a Civil Engineering degree, 2+ years relevant experience, and strong written E
Design original, research-style computational math problems and provide reproducible Python solutions for OpenTrain AI; 20+ hrs/week, contractor, $15–$60/hr. Requires a math degree, 2+ years relevant experience, and strong Python/NumPy/SciPy/SymPy skills.
Join OpenTrain AI as a part-time Senior Python Code Reviewer to audit AI-generated Python snippets in containerized sandboxes; $18/hr, remote, under 20 hrs/week. Bring 7+ years of Python experience and mandatory Docker proficiency to validate correctness, security, and guideline compliance.