AI Chatbot Workflow Testing & Evaluation (Independent Research Project)
Evaluated AI chatbot outputs across multiple LLMs by designing structured test cases and assessing conversational quality. Documented recurring failure patterns such as hallucinations, instruction drift, and context loss with severity ratings. Applied a repeatable multi-criterion rubric to consistently score prompt-response pairs at scale. • Built 40+ structured test cases for coherence, factual accuracy, and task-completion quality • Produced actionable evaluation reports with severity ratings for model-specific failures • Used a 5-criterion rubric (accuracy, relevance, fluency, safety, usefulness) across 100+ evaluations • Benchmarked ChatGPT (GPT-4), Claude 3, and Gemini 1.5 Pro