Senior AI Trainer & Data Annotation Specialist
Created realistic, user-like prompts across general domains and evaluated model instruction-following behavior and output quality. Reviewed and edited AI-generated responses for clarity, factual correctness, tone, and usefulness to improve overall response quality. Authored human reference answers and prompt variations to support supervised fine-tuning and evaluation benchmarks. • Produced prompt variations and controlled test sets to measure nuanced output differences and edge-case behavior. • Ranked and compared multiple model responses under length/style constraints and task-specific instructions using preference-style labeling. • Identified failure cases such as hallucinations, ambiguity, and safety issues and documented findings for QA updates. • Supported evaluation workflows by generating high-quality reference content aligned with selected prompts.