Outlier
I first determine if a STEM-related prompt can be evaluated. Once verified, I am presented two responses from different LLMs and I choose which response is better. I provide a detailed justification for my decision in each case.
Hire this AI Trainer
Sign in or create an account to invite AI Trainers to your job.
I've been actively working as a freelance AI trainer and response evaluator on Outlier focused primarily on STEM-heavy domains: mathematics, physics, programming in C++, and formal logic and puzzle problems. My standard workflow is to independently solve each problem to establish ground truth — often writing Python scripts for computational verification — before scoring model responses against a structured rubric covering instruction following, accuracy, verbosity, format, and writing style, followed by a winner verdict and itemized error flags. This ground-truth-first approach reliably catches subtle reasoning failures, fabricated citations, and silently incorrect derivations that surface-level review misses.
I first determine if a STEM-related prompt can be evaluated. Once verified, I am presented two responses from different LLMs and I choose which response is better. I provide a detailed justification for my decision in each case.
Master's, Physics
Bachelor's, Physics
Graduate Research Assistant
Graduate Research Assistant