Modern Work Architect — Copilot Agents and AI output evaluation discipline (prompt/rubric/validation)
Evaluated LLM outputs for technical accuracy by flagging hallucinations, outdated guidance, and incorrect configurations against production knowledge. Used rubric-based side-by-side ranking across instruction-following, factual accuracy, code correctness, reasoning quality, and tone, producing detailed rationale writeups. Ensured evaluation outputs were validated before reaching enterprise production settings and aligned with deployment governance requirements. • Ranked model responses against explicit criteria for safety, correctness, and reasoning • Authored reference/gold responses and rubric explanations for specialized technical domains • Designed prompt boundaries and grounding approaches that affect evaluation outcomes • Performed pre-production validation of automated/AI tooling outputs