Privacy-First Local AI Agent (GenAI & MLOps) — human + LLM-as-a-judge evaluation
Conducted human and LLM-as-a-judge evaluations to assess reasoning quality and retrieval-augmented generation outputs. Evaluated Chain-of-Thought (CoT) reasoning and RAG accuracy against predefined criteria and guardrails. Used these evaluations to identify failure modes and inform iterative improvements to the local agent workflow. • Assessed CoT reasoning quality using human review and automated judging • Checked RAG accuracy using retrieved context from a local vector store • Applied prompt and guardrail constraints to standardize evaluation • Reported and refined results through an iterative model engineering loop