AI Data Labeling & Prompt Evaluation Specialist
Evaluate and rate AI-generated responses across dimensions including factual accuracy, instruction following, helpfulness, coherence, and stylistic. Maintain consistent quality scores across batches, progressing through platform tier levels based on performance. Assess AI-generated outputs for quality, safety, and alignment with user intent across diverse task categories. Complete structured annotation tasks requiring careful judgment on model behavior, tone, and factual grounding. Flag problematic outputs including hallucinations, unsafe content, and instruction non-compliance.