LLM Response Quality Annotation & Evaluation — Independent Research
Built a 5-dimension LLM response quality annotation rubric covering factual accuracy, reasoning, instruction-following, coherence, and safety. Annotated 500+ prompt–response pairs to evaluate LLM outputs against the rubric dimensions. Achieved Cohen’s kappa above 0.82 across all dimensions, indicating strong labeling consistency. • Defined rubric criteria for multi-dimensional LLM output evaluation • Performed prompt–response annotation aligned to quality dimensions • Conducted/managed rater agreement scoring using Cohen’s κ • Evaluated LLM responses using structured rubric-based judgments