Generalist Evaluator with Mercor
Reviewed and evaluated LLM responses through Mercor, focusing on user intent alignment, instruction-following, accuracy, clarity, and response quality. Tasks included identifying when AI assistants misunderstood user intent, added repetitive filler, or gave ambiguous answers. I provided structured ratings and written feedback to help improve model performance and output reliability.