Generalist/Aether AI research evaluator (visual and multimodal content)
Evaluated AI-generated multimodal outputs using structured evaluation guidelines and consistent scoring rubrics. Checked outputs for factual and guideline adherence, reporting inaccuracies and quality issues for research and production workflows. Communicated clear quality assessments and improvement notes to enable model iteration. • Used detailed prompts and rubrics for consistent assessment • Flagged hallucinations and coherence problems • Delivered written feedback supporting research objectives • Met strict productivity and accuracy thresholds independently