Data Processing Engineer - Data Labeling Project
Led the end-to-end workflow of data cleaning and annotation for over 15,000 structured and unstructured text samples. Developed standardized labeling protocols to unify 12 label categories with consistency exceeding 95%. Performed dataset partitioning and annotation to support iterative NLP model training cycles. • Utilized Python and libraries like Pandas, Regex, and OpenPyXL to automate data deduplication and transformation. • Improved data usability by 68% and boosted processing efficiency by 80% over manual approaches. • Independently handled training, validation, and test set annotation tasks for five NLP model versions. • Contributed directly to a 12.5% average increase in model F1 scores through accurate and consistent labeling.