Research Assistant at Relational Cognition Lab, UCI (January 2026 – Present) building LLM evaluation harnesses and experiment orchestration.
Built an evaluation harness to probe and compare model representations across Gemma, Qwen, and Llama using standardized prompt sampling and aggregation. Implemented a noise-estimation baseline by injecting random text into prompts to quantify variance for defensible comparisons. Orchestrated repeatable experiments with an evaluation framework and tracking to ensure reproducibility, then deployed an inference service to serve embeddings for downstream analysis. • Standardized prompt sampling and aggregation for multi-model evaluation • Noise-injection baseline for variance/robustness measurement • Experiment orchestration with reproducible tracking and metrics • FastAPI + Redis embedding inference service to reduce downstream analysis latency