Freelance Content Evaluator / AI Trainer (Remote)
Evaluated and graded AI-generated responses using Helpfulness, Honesty, and Harmlessness criteria. Performed prompt engineering to probe model capability boundaries and identify failure modes. Detected and corrected hallucinations and ensured outputs met accuracy, safety, and guideline requirements. • LLM response evaluation and rubric scoring • Hallucination detection and correction of fact errors • Follow evolving labeling/project guidelines for data quality • Team communication and deadline-driven review cycles