Senior Software Engineer / Technical Lead — AI Platform, Full-Stack Systems & Cloud Infrastructure (AI training/evaluation work in production)
Built and maintained LLM verification and evaluation frameworks, including golden datasets and evaluation harnesses, to measure and improve AI-generated output quality at scale. Implemented hallucination reduction techniques and verification loops using OpenAI API and LLM tooling to ensure correct, secure, and reliable results before deployment. Established validation standards and quality gates for AI-generated code and outputs across production workflows. • Used golden datasets, eval harnesses, and accuracy/reliability metrics to calibrate model behavior • Implemented verification loops and guardrails for hallucination reduction and secure outputs • Created standardized contributor evaluation rubrics and calibration mechanisms • Supported AI-native development workflows by enforcing validation standards for AI-generated artifacts