Co-Founder & AI Product Lead — Orceum
Led the design of LLM response validation pipelines and evaluation rubrics to reduce hallucinations and improve context retention for vertical AI agents. This work functioned as AI labeling via structured scoring and validation of model outputs against quality criteria. • Built and iterated LLM response validation pipelines • Developed model evaluation rubrics for prototype scaling • Benchmarked and improved response correctness for vertical agents • Coordinated product and engineering delivery around evaluation quality