AI Training & Content Evaluator (Freelance)
Evaluated complex text datasets and AI-generated responses using criteria focused on linguistic precision, factual correctness, truthfulness, harmlessness, and relevance. Performed ranking and quality assessment to identify better responses for downstream use. Flagged edge cases and potential failure points to improve model robustness and reliability.• Evaluated AI outputs for accuracy, safety, tone, and helpfulness.• Analyzed and ranked multiple responses against truthfulness and relevance criteria.• Identified and flagged edge cases in data samples.• Collaborated on research-heavy review of technical and creative content accuracy.