LLM Evaluator
Project Description: Evaluated and optimized a Large Language Model (LLM) to improve conversational accuracy, reasoning capabilities, and factual truthfulness. Authored complex prompts to test model boundaries and applied Reinforcement Learning from Human Feedback (RLHF) to score multi-turn chatbot responses. Reviewed model outputs for grammar, tone, and logical consistency, successfully reducing model hallucinations and improving response quality metrics.