LLM Evaluation Dashboard (Project)
Created an LLM evaluation dashboard to compare model outputs across multiple benchmarks for data-driven model selection. Aggregated and visualized evaluation metrics using Pandas and NumPy to support benchmark analysis. Integrated an external evaluation API to test financial signal performance on quantitative reasoning tasks. • Performed LLM output evaluation across benchmarks. • Used Pandas/NumPy for metric aggregation and visualization prep. • Integrated WorldQuant Brain API for domain-specific evaluation. • Presented interactive evaluation results via a React/Node.js frontend.