AI Engineer & Data Annotation Designer
No description provided.
Hire this AI Trainer
Sign in or create an account to invite AI Trainers to your job.
AI Engineer & Data Annotation Designer (Text/NLP annotation pipelines). Brings 4+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Label Studio, Prodigy, and Other. Education includes Bachelor of Engineering, Hunan University (2020). AI-training focus includes data types such as Text, Image, and Audio and labeling workflows including Entity (NER), Segmentation, and Classification.
No description provided.
Architected end-to-end multimodal data annotation pipelines including video modalities for AI training. Built and maintained annotation standards, taxonomies, and quality rubrics to reduce inter-annotator disagreement by 32% while meeting throughput of 50,000+ labeled samples weekly. Implemented tiered review processes that delivered 98.5% final label accuracy on production datasets. • Led 15+ annotators across multi-stage QA workflows to ensure label consistency. • Developed LLM-based AI agents for automated data preprocessing, including intelligent cleaning, entity extraction, and semantic deduplication, reducing manual work by 40%. • Created multi-agent RAG tool-use systems for autonomous data curation such as cross-source alignment and conflict resolution. • Oversaw dataset versioning and documentation using Label Studio, Prodigy, and internal tooling.
Designed multimodal AI training dataset annotation pipelines including audio modalities to support large-scale model training across the organization. Developed annotation guidelines, taxonomies, and quality rubrics that decreased inter-annotator disagreement by 32% while maintaining 50,000+ labeled samples per week. Established tiered review workflows to achieve 98.5% final label accuracy on production datasets. • Led annotation team execution through initial labeling, peer audit, and expert adjudication. • Applied LLM agent workflows for automated preprocessing such as data cleaning and semantic deduplication, cutting manual preprocessing by 40%. • Collaborated with ML teams on dataset composition iteration using error analysis and bias audits improving F1 by 6–12 points. • Managed full data lifecycle from ingestion and schema design to versioned dataset releases.
Architected multimodal annotation pipelines that included image labeling for AI training datasets, supporting 5+ concurrent model training initiatives. Created comprehensive annotation guidelines, taxonomies, and quality rubrics that reduced inter-annotator disagreement by 32% while meeting throughput targets of 50,000+ labeled samples per week. Coordinated multi-level QA review stages to reach 98.5% final label accuracy on production datasets. • Led cross-functional annotation teams of 15+ annotators with initial labeling, peer audit, and expert adjudication. • Implemented automated preprocessing with LLM-based agents (data cleaning, entity extraction, and semantic deduplication) reducing manual effort by 40%. • Built multi-agent RAG and tool-use frameworks for autonomous data curation such as cross-source alignment and conflict resolution. • Used Label Studio, Prodigy, and internal platforms to manage labeling operations and dataset versions.
Led end-to-end annotation pipeline design for text-centric multimodal datasets to support production AI model training across 5+ concurrent initiatives. Built annotation guidelines, taxonomies, and quality rubrics to reduce inter-annotator disagreement by 32% while sustaining 50,000+ labeled samples per week. Implemented tiered review workflows (initial labeling → peer audit → expert adjudication) achieving 98.5% final label accuracy on production datasets. • Authored SOPs and technical documentation for annotation workflows and reproducibility. • Applied error analysis and data bias audits with ML research teams to improve model F1 by 6–12 points. • Developed LLM prompt optimization and few-shot templates to speed up annotation on text classification and NER tasks by 25%. • Managed dataset schema design and versioned dataset releases across the data lifecycle.
Bachelor of Engineering, Environment Engineering
AI Engineer & Data Annotation Designer
Java Developer