AI Language Model Evaluator & Trainer — Outlier AI (August 2025 – Present)
Evaluates LLM-generated outputs for factual accuracy, logical consistency, and contextual relevance across multiple subject domains. Refines prompts to improve response clarity, specificity, and alignment with diverse user intents while supporting standardized calibration with evaluation teams. Conducts systematic web-based fact-checking of AI-generated claims using authoritative sources and applies output validation and quality scoring to maintain benchmark standards. • Labeling includes accuracy, coherence, and contextual relevance judgments • Applies response quality scoring and validation/rubrics • Performs web-based claim verification • Collaborates to standardize annotation workflows and calibration criteria