Independent Contributor – AI & Content Evaluation
Evaluated AI-generated written content for factual accuracy, coherence, tone, and adherence to prompt instructions using multi-criteria rubrics. Detected hallucinations, reasoning errors, and guideline violations in LLM outputs and produced structured written feedback to support model improvement. Performed quality assessment work intended to raise helpfulness, truthfulness, and harmlessness metrics across response dimensions. • Rated responses using rubric-based scoring across helpfulness, truthfulness, and harmlessness. • Provided structured feedback for model improvement, including identified failure modes. • Assessed compliance with prompt instructions and safety-related guideline adherence. • Performed critical checking for accuracy, coherence, and other evaluation criteria.