Technical AI Content Specialist (RLHF)
Evaluated Side-by-Side (SBS) and ranking outputs using the HHH rubric (Helpfulness, Honesty, Harmlessness) to improve model behavior. Performed hallucination detection and technical fact checking for Python/STEM prompts, using ground-truth reasoning to assess logic and adherence to instructions. Provided structured feedback to strengthen reasoning quality and reduce incorrect or unsafe outputs. • Helpfulness/Honesty/Harmlessness scoring and comparative judgments (SBS) • Hallucination detection and technical fact checking for STEM prompts • Instruction-following and reasoning adherence evaluation • Feedback grounded in correct reasoning and auditing of responses