AI Trainer & Evaluation Specialist (Freelance, Remote)
Evaluated hundreds of AI-generated outputs for factual precision, logical flow, tone compliance, and adherence to multi-turn prompt instructions. Flagged edge-case anomalies including formatting deviations, structural errors, and subtle hallucinated or unreliable data across multiple domains. Produced justification-style reports to communicate model performance, linguistic flaws, and factual gaps to engineering teams. • Reviewed LLM response quality against provided guidelines and prompt constraints. • Identified and escalated anomalies, hallucinations, and taxonomy misalignment. • Authored audit-ready justification reports for downstream improvements. • Applied evolving evaluation taxonomy frameworks across high-volume queues.