AI Evaluation & Rubric Specialist
As an AI Evaluation & Rubric Specialist, you evaluated and ranked AI-generated customer service responses using quality, tone, reasoning, professionalism, and policy alignment criteria. You created detailed rubrics and evaluation criteria to assess LLM performance in simulated business and customer support scenarios. You built personas and digital environments and ran AI agents against structured scenarios, then analyzed weaknesses and provided structured feedback to guide performance improvements. • Evaluated multiple AI responses and produced rankings based on professionalism, human-like communication, reasoning quality, tone, and policy alignment. • Built detailed rubrics and evaluation criteria for LLM assessment across simulated scenarios. • Designed realistic tasks and digital environments using tools such as Gmail, Slack, Calendar, WhatsApp, and Google Drive. • Ran agents on structured prompts, analyzed outputs, and supported improvements via feedback and hinting systems.