Create advanced STEM questions, score AI-generated solutions, and build detailed rubrics for model training. This remote US contractor role pays $100 per hour and requires 20+ hours weekly.
The Work
You will help assess and improve advanced AI systems by turning expert technical judgment into clear questions, scoring rules, and reference answers. The work focuses on STEM topics and requires careful review of both final answers and the reasoning behind them.
- Write expert-level technical questions that test genuine, multi-step reasoning and reveal weaknesses in AI systems.
- Create detailed rubrics covering required reasoning, correct methods, acceptable alternatives, partial credit, and critical errors.
- Review AI-generated solutions for correctness, reasoning quality, completeness, methodology, and technical precision.
- Find hidden reasoning errors, including invalid assumptions or flawed logic behind a correct final answer.
- Review other experts’ rubrics for technical accuracy, clarity, completeness, ambiguity, redundancy, and consistent scoring.
- Create expert answers, explanations, and gold-standard evaluations for reinforcement learning and supervised fine-tuning workflows.
What It Pays and Takes
This is a remote, part-time contractor role for candidates based in the United States. The posting lists the role as entry level, while the requirements call for substantial technical expertise and at least three years of relevant experience.
- Pay: $100 per hour.
- Time: 20 or more hours per week.
- Location: United States.
- Language: English at professional C1 level or higher.
- Education: Bachelor’s degree or higher in mathematics, physics, chemistry, engineering, computer science, statistics, applied science, or another technical field.
- Experience: At least three years of professional, academic, research, or industry experience in your technical domain.
- Expertise: Advanced knowledge in a clearly defined technical field or specialization.
- You must be able to write precise technical explanations and distinguish valid reasoning from flawed reasoning.
- Experience with exam writing, grading, peer review, research review, technical quality assurance, standards development, or assessment design is preferred.
- Experience with AI evaluation, reinforcement learning from human feedback, supervised fine-tuning, benchmarking, data annotation, prompt design, or large language model evaluation is preferred.
How It Works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI Training Work
AI training is the human work behind systems that learn from examples, including writing questions, rating model responses, and checking reasoning. OpenTrain helps people find and build careers in this field, where strong technical judgment is used to make AI systems more accurate and reliable.