Data Analyst and AI Evaluation Specialist - Filo
In this role, you analyzed structured AI interaction records to identify behavioral patterns, failure modes, and non-stationary performance trends. You designed classification frameworks and failure taxonomies to support consistent evaluation across hallucination, reasoning breakdown, instruction drift, and multi-step task failures. You built Python-based workflows for data quality assessment, model monitoring, and reproducible reporting using common data and ML tooling. • Analyzed 3,000+ structured AI interaction records for performance trends and behavioral drift • Built classification frameworks and failure taxonomies for common model failure modes • Developed Python workflows for structured analysis, data quality assessment, and monitoring • Produced reproducible reporting to support AI evaluation and insights