Machine Learning Intern, Production Runtime Experience (Goldman Sachs)
Benchmarked multiple LLM models within a multi-agent LangGraph architecture to improve agent performance and reduce unnecessary or low-quality outputs. • Engineered a Python agentic framework for deep analysis and discovery over a very large graph dataset. • Designed a LangGraph multi-agent system with state isolation and guardrails to minimize errors and content bloat. • Ran and recorded evaluation across six LLM models per agent (runtime, responses, and token costs) for data-driven model selection. • Integrated agent tools with existing services to support production workflow execution.