AI Benchmark Evaluator and Reviewer
Evaluated and reviewed structured AI benchmark items by assessing model reasoning, factual accuracy, instruction-following, and edge-case handling. Applied rubric-based scoring and consistency checks to identify ambiguous answers, missing correct options, and cases where multiple choices were simultaneously defensible. Conducted multi-LLM adjudication by distinguishing hallucinations from genuine factual ambiguity using independent source verification. • Produced 300+ benchmark items and completed 1,000+ structured AI output evaluations. • Corrected and re-authored evaluation rubrics to be clearer, more atomic, and self-contained. • Detected errors missed by question authors through strict rubric application. • Adjudicated conflicting outputs across large language models with verified evidence.