LLM Evaluator / AI Trainer – Multi-Agent Cloud Resource Scheduler Project
Led the evaluation and structured validation of LLM agentic outputs for planning and scheduling tasks in a multi-agent cloud resource management project. Focused on schema-constrained tool-call outputs, parse-success monitoring, hallucination detection, and deterministic fallback implementations. Evaluated model reasoning by transforming server state vectors into natural-language prompts and measuring compliance with hard constraints. • Designed and implemented JSON schema-based tool-call evaluations for LLM agents. • Benchmarked LLM outputs against heuristic and rule-based baselines. • Created self-distilled SFT data from decision traces for fine-tuning. • Built reproducible test pipelines to ensure labeling and evaluation quality.