Data Scientist (Labeling & QA of training datasets) at Google (2018 to Present)
You collaborated with ML teams to label and QA large-scale text datasets used for recommendation and churn prediction. You defined annotation guidelines and managed vendor labeling pipelines through Labelbox to ensure consistent dataset quality. You integrated active learning feedback loops to connect model predictions with labeling priorities and improve iterative labeling efficiency. • Defined labeling guidelines and taxonomy requirements with data ops teams • Performed QA and inter-annotator quality checks for labeled datasets • Coordinated vendor labeling workflow execution in Labelbox • Used active learning to reprioritize data for annotation based on model uncertainty.