A.I. Evaluation and Training, Alphabet Inc. (second stint)
Assessed subjective quality signals for AI responses, scoring conversational metrics such as tone, helpfulness, clarity, and user satisfaction. Verified whether outputs adhered to policy and safety constraints by checking for bias, toxic undertones, policy violations, and subtle hallucinations. Tracked and validated the logical steps an AI agent took to ensure the final answer was grounded in facts rather than guesswork. • Quality scoring of conversational responses • Policy/safety boundary checks (bias, toxicity, violations) • Hallucination detection and edge-case spotting • Reasoning-step validation of agent outputs