Over the past year, I have gained hands-on experience as an AI Training Specialist and Data Labeler, specializing in aud
Over the past year, I have gained hands-on experience as an AI Training Specialist and Data Labeler, specializing in audio transcription and speech data annotation to optimize machine learning models. Working remotely on large-scale annotation pipelines—including Appen and CrowdGen initiatives like Project Jigglypuff—my core responsibilities involved reviewing machine-generated data, correcting timestamp segment boundaries, and modifying text outputs to reflect spoken audio perfectly. I strictly adhered to complex acoustic formatting guidelines, handling linguistic gray areas such as distinguishing silence from natural pauses, managing code-switching, and ensuring precise Word Error Rate (WER) metrics. A major focus of my workflow involved complex multi-speaker segmentation and localized environmental tagging. I accurately isolated primary speakers from background noises, applying precise formatting tags like <nonprimaryspeakertalking> for overlapping dialogue or environmental sounds (such as [cough], [laugh], or [music]). Managing large volumes of hourly data units across strict quality control thresholds has given me a deep understanding of data pre-processing, metadata tracking, and quality assurance workflows within the AI development lifecycle.