AI Data Specialist, Project Diamond (Remote)
Rigorously evaluated and annotated generative LLM outputs to improve downstream accuracy and safety, focusing on logical consistency, tone, and bias reduction. • Labeled multi-turn model behavior against project-specific helpfulness guidelines. • Authored and used complex prompt variations to identify edge cases and hallucination triggers. • Translated ambiguous queries into structured evaluation metrics to enforce strict alignment with rating criteria. • Performed ongoing stress testing and documentation of failure modes for safer model outputs.