AI Evaluation Specialist – Outlier AI
Evaluated AI-generated outputs across writing, reasoning, and instruction-following tasks to assess quality and model behavior. Reviewed responses for truthfulness, localization, conciseness, and harmful content while ensuring consistency with evaluation criteria. Supplied detailed feedback to improve model performance and response quality. • Assessed instruction-following, reasoning quality, and writing outputs • Flagged issues related to truthfulness and harmful content • Provided structured, actionable feedback for improvements • Handled large volumes while maintaining accuracy