AI Evaluation Engineer (Task Creator, QC Reviewer & Task Rework Specialist)
Served as an AI evaluation engineer responsible for creating and correcting graded evaluation tasks for autonomous AI agent benchmarking. Tuned scoring pipelines using live rollout data, improving grader logic and prompt alignment to rework underperforming evaluation tasks. Performed QC review using a structured 20-point framework to validate grader fairness, determinism, and overall evaluation quality. • Reworked evaluation tasks after identifying grader failures (e.g., shell escaping issues, kubeconfig path errors, rollout faults). • Developed graded scenarios for broken Kubernetes clusters and misconfigured DevOps components (KEDA, HPA, Kong Gateway, Loki, Redis TLS, StatefulSets). • Progressed from Task Creator to QC Reviewer and applied the QC framework to enhance benchmark reliability. • Ranked in the top 10 among 500+ contributors on the Horizon AI benchmarking platform.