Senior Data Scientist, AI Evaluation
As a Senior Data Scientist at Google, I designed and executed structured evaluation workflows for AI-generated content. I focused on reasoning consistency, factual reliability, and benchmark validation for large-scale LLM applications. My responsibilities included collaborating on AI benchmarking and enhancing data quality standards in AI evaluation contexts. • Developed testing frameworks for evaluating LLM model outputs. • Implemented benchmark validation tasks to ensure high factual reliability. • Led efforts in model reliability and AI-generated response quality control. • Improved cross-team processes for human-in-the-loop and analytical quality assurance.