Python-based LLM evaluation framework for customer feedback classification/summarization and LLM reliability/latency benchmarking
Led development of an LLM evaluation framework to benchmark model outputs for customer-feedback classification and summarization tasks. Created reusable model-provider architecture with validated structured outputs to assess quality and reliability. Built automated evaluation pipelines comparing model accuracy, latency, and output reliability across multiple LLM families. • Benchmarked customer feedback classification • Implemented provider support for Ollama and OpenAI • Added structured output validation via Pydantic • Automated comparison of accuracy/latency/reliability