AI Product Engineer — Midas AI (multi-agent agentic research platform with LLM evaluation/grounding)
Built and evaluated an LLM-driven multi-agent research pipeline for financial equity research using LangGraph and hybrid RAG sources. Implemented heuristic + LLM-as-judge scoring to verify report claims against ground-truth evidence and enforce citation-level provenance. Designed quality gating across retrieval, planner/coordinator, research, and writer stages to improve factuality and reduce hallucinations. • Scored reports on 8 weighted dimensions using LLM-as-judge • Grounded outputs by checking each claim against filings, financial APIs, and live web evidence • Implemented per-claim citation provenance with deterministic computation and source-priority fusion • Added anti-hallucination layers and reproducible evaluation for citation-level accuracy