Senior AI Output Reviewer & Quality Rater
This role involved evaluating and ranking AI-generated responses for accuracy, coherence, and safety. Key tasks included identifying hallucinations, logical errors, and prompt vulnerabilities in both code and text outputs. The position required detailed justification writing and consistent adherence to complex annotation guidelines. • Ranked and scored LLM outputs based on truthfulness and helpfulness. • Flagged hallucinations, prompt injections, and logical inconsistencies. • Reviewed and corrected Python/SQL code snippets for SFT baseline data. • Maintained a 98%+ accuracy rate following 50+ page guidelines.