Review AI-generated astronomy and space-science responses for scientific accuracy, assumptions, units, and uncertainty; remote contract, 20+ hrs/week, with pay from $40–$100/hr depending on location and qualification.
Generative AI & RLHF
Remote Hourly · $40–$100/hr
$40–$100/hr
Compensation
17 countries
Eligibility
Intermediate
Experience
Jul 8, 2026
Posted
Open to applicants in
Bangladesh Bhutan Brazil Cambodia Germany India Indonesia Malaysia Nepal Pakistan Singapore Sri Lanka Thailand Philippines United States Timor-Leste Vietnam
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform for people who build careers in AI training and data labeling. We connect qualified contributors with specialist projects, help you consolidate work and build a single AI-training portfolio, and hire contractors directly for focused evaluation and annotation work.
Creating an OpenTrain account is free; through OpenTrain you can find remote, flexible projects that let you apply domain expertise to shape how AI systems behave.
About AI training work
AI training (data labeling / annotation / human feedback) is the human side of building intelligent systems: people create, check, and judge examples that models learn from. This role focuses on evaluating technical model outputs—an opportunity to influence how scientific AI explains reasoning, states uncertainty, and avoids overconfident or incorrect claims.
Work is 100% remote and highly flexible; many contributors treat it as part-time or freelance work.
Specialist domain knowledge is rewarded; technical reviewers directly shape model behavior in their fields.
The role
OpenTrain is recruiting Astronomy & Space Science AI Evaluators to assess AI-generated responses, prompts, and step-by-step solutions across astronomy and space science. You will judge scientific accuracy, logical soundness, completeness, units, assumptions, approximations, uncertainty, and limits of inference, and provide structured feedback that improves model outputs.
Commitment: 20+ hours per week (remote, contract, part-time).
Pay: advertised hourly range $40–$100 USD; final rate varies by country and experience (input lists hourly up to 40, 60, or 100 depending on country).
Labeling focus: evaluation ratings and RLHF-style judgments on TEXT outputs.
What you'll do
Evaluate AI-generated answers for scientific accuracy, logical soundness, and completeness.
Review and refine prompts, model answers, and step-by-step solutions in astronomy and space science.
Check units, assumptions, approximations, uncertainty estimates, and limits of inference.
Identify conceptual errors, bad assumptions, missing constraints, and formula mistakes.
Assess outputs across orbital mechanics, stellar and galactic astrophysics, cosmology, planetary science, space physics, and data interpretation.
Design and apply realistic, research-style challenges and calculations to probe model reasoning.
Provide clear, structured feedback that explains errors, suggests corrections, and advises on communicating uncertainty.
Requirements
You must have a demonstrated background in astronomy or a closely related space science discipline and be able to perform quantitative sanity checks, unit checks, and reasoning about uncertainty.
Minimum: Bachelor’s degree in physics, astronomy, space science, engineering, or a closely related field; Master’s or PhD strongly preferred.
Experience: 4+ years of professional experience in astronomy or a closely related space science domain.
Deep knowledge of core fundamentals such as orbital mechanics, radiative processes, stellar and galactic physics, coordinate systems, time standards, statistics, and scientific modeling assumptions.
Strong quantitative reasoning: validate results using orders of magnitude, unit consistency, limiting cases, and error propagation.
Ability to communicate model limitations responsibly and challenge overconfident scientific claims.
Nice-to-have technical skills
Experience with AI data training, annotation, red-teaming, or evaluating AI-generated technical content.
Familiarity with Python, NumPy, SciPy, data pipelines, catalog queries, photometry, spectroscopy, or mission datasets.
Background in orbit determination, photometry, spectroscopy, or other quantitative analysis workflows.
Work arrangement and other details
This is a contractor engagement managed and hired by OpenTrain. We welcome qualified contributors in multiple countries (listed in the posting) and expect remote collaboration.
The role supports high-quality scientific evaluation of advanced model outputs; reviewers must be comfortable providing structured, technical feedback and working with scientific assumptions and uncertainties.
Employment types: Contractor, Part-time.
Languages: English required.
Countries: VN, TL, TH, SG, PK, PH, NP, MY, LK, KH, ID, BT, BD, IN, BR, DE, US (Open to qualified contributors in these countries).
How to apply / next steps
Create or sign in to your free OpenTrain account, complete your profile with relevant astronomy and technical experience, and submit an application for this evaluator role. Selected contributors will receive contractor engagement details and onboarding for the specific evaluation tasks.
When you apply, highlight relevant publications, mission work, quantitative analyses, or prior AI-evaluation experience to help match you to advanced scientific review tasks.
Lead QA for astronomy and astrophysics AI training: review model outputs and trainer submissions for scientific, mathematical, and editorial accuracy. Remote US-only contractor role, 20+ hours/week, up to $110/hour.
Join OpenTrain to evaluate frontier physics research and judge AI model reasoning, earning $80–$110/hr. This part-time, remote US role requires a strong publication record and active research experience (PhD/postdoc preferred).
Join OpenTrain AI to evaluate and improve physics-focused AI outputs—paid contract work at $80/hr, part-time (minimum ~17–20 hrs/week). Use your physics degree to spot subtle errors, write step-by-step solutions, and rate model responses with detailed rubrics.