LLM Response Evaluation
Evaluated AI-generated responses for accuracy, relevance, completeness, instruction following, and factual consistency. Compared multiple LLM outputs using structured evaluation criteria, identified hallucinations and reasoning errors, ranked responses based on quality, and documented feedback to improve model performance. Worked with prompt-response datasets and followed annotation guidelines to ensure consistent evaluations.