Senior Machine Learning Engineer & Data Scientist — DoorDash (Dec 2022–Present) focused on LLM evaluation and agentic AI workflow evaluation
Built and deployed LLM-as-a-Judge evaluation frameworks to detect hallucinations, validate semantic correctness, and benchmark model performance before release. Performed offline/online evaluation of LLM-powered semantic search and RAG systems to support safe, observable deployments. Coordinated evaluation and release readiness across multi-agent workflows that enforce structured outputs and grounded generation. • Hallucination detection and semantic correctness validation for production LLM releases • Benchmarking multiple models (e.g., GPT-4/Claude) for quality/latency/token-efficiency tradeoffs • Offline/online evaluation of RAG pipelines using embedding-based retrieval • Structured validation layers and guardrails integrated into evaluation and agent execution