Physics Expert
I work on RLHF (reinforcement learning from human feedback) to improve frontier large language models on physics, from undergraduate through PhD level. My work spans authoring headroom questions, including frontier-level and image-based items designed to challenge state-of-the-art models and surface where their reasoning breaks down, as well as CUJ (Critical User Journey) tasks. I build rubrics for both text- and image-based physics questions, and I assess LLM outputs across modalities, including text, image, HTML, and video. The core of the role is rigorous, step-level evaluation of physics reasoning: grading not just final-answer correctness but the validity of each step, and identifying solutions that appear correct while concealing a subtle conceptual or mathematical error.