AI Evaluator & Data Annotator, Product Design & UX Expert (Mercor AI / OpenAI)
Evaluated AI-generated UI and conversational AI outputs for a frontier AI lab using structured rubrics. Rated design screenshots on multiple visual and usability dimensions and provided 2–3 sentence written justifications per rating. Performed comparative ranking and consistency checks to align scores with best-to-worst order. • Rate UI designs on a 1–7 Likert scale across hierarchy, typography, color, components, navigation, and aesthetic. • Rank/reorder design screenshots best-to-worst and flag inconsistencies between ratings and order. • Evaluate combined conversational AI responses with generative UI for user progression using cognitive load, intent adaptation, and tone criteria. • Assess front-end code quality and instruction adherence on a 1–5 scale and escalate PII/policy violations.