AI Evaluator / RLHF & Safety Reviewer (AI-generated response quality and preference evaluation)
Performed comparative AI response evaluation to support RLHF-style human preference learning and iterative model improvement. Assessed prompt-response quality using rubric-based judgments for safety, policy alignment, hallucination likelihood, and conversational calibration. Produced structured, evidence-based reviewer notes and rankings to reflect human-like preferences and detect subtle reasoning issues. • Response A vs B comparative ranking and preference scoring. • Safety, privacy, and governance reasoning checks including hallucination/misinformation detection. • Emotional calibration and tone evaluation for conversational quality. • Ambiguity and tradeoff analysis with escalation and justification writing.