Senior AI Evaluator / RLHF Data Specialist - InsightLabel AI
Led LLM evaluation efforts using structured rubrics to measure quality, reasoning, and conversational performance. Performed RLHF preference ranking and scoring to support human-feedback optimization of model behavior. Designed and maintained annotation and testing processes to ensure dataset consistency, including adversarial evaluation for hallucinations and bias. • Create evaluation rubrics and scoring criteria for conversational AI • Execute RLHF ranking and preference scoring workflows • Develop dataset guidelines for consistency and quality assurance • Run adversarial tests to detect hallucinations and bias