AI Model Training Specialist at Outlier (AI training/annotation tooling and labeled dataset preparation)
Trained and improved AI model datasets by leveraging structured data extraction and labeling workflows using Python and Pandas, contributing to higher NLG accuracy. Collaborated with developers and stakeholders to redesign annotation tools and automate repetitive tagging to increase labeling throughput. Performed data quality audits and statistical validation to ensure labeled datasets were balanced and reduced downstream bias in model outputs. • Extracted and prepared structured training data for model learning • Automated repetitive tagging and reduced manual processing time • Conducted automated data quality audits on labeled datasets • Validated training distribution assumptions using statistical tests