RLHF and Output Quality Evaluation for a Production LLM Assistant
Designed and ran a structured evaluation process for the AI generated responses powering a real time LLM assistant.Across three task modes (conversational, behavioral, and technical), I rated and ranked model outputs for factual accuracy, relevance, tone, and adherence to instructions, and authored explicit acceptance and rejection criteria that functioned as annotation guidelines