AI Model Evaluator – LLM & Agent Systems (Contract/Independent)
Conducted structured evaluation of LLM and agent outputs using predefined guidelines and scoring rubrics to verify quality, relevance, and logical consistency. Reviewed multi-step autonomous agent workflows, reasoning traces, and output artifacts to assess accuracy and completeness. Identified recurring failure modes such as hallucinations, incomplete reasoning, and instruction misalignment, then documented findings to support model improvement. • Evaluated prompt-response quality against rubric criteria • Calibrated with other reviewers to maintain consistent evaluation standards • Produced structured written feedback for AI benchmarking and refinement • Assessed compliance with defined standards and benchmarking objectives