African Voice Project
This project involved linguistic data annotation for Natural Language processing (NLP) research focusing on Yoruba and English data. I worked on transcribing and segmenting audio recordings, including speaker identification and timestamp alignment using ELAN. The main goal was to prepare high-quality annotated datasets suitable for machine learning and language technology applications. My tasks included labelling spoken data, ensuring accurate segmentation of speech units, and maintaining consistency in annotation across datasets. I ensured quality through careful observation and review of transcriptions, cross-checking for accuracy in speaker labelling, and maintaining alignment precision in timestamps. The project demanded strong attention to detail, consistency, and an understanding of linguistic structure to ensure the datasets were suitable for NLP model training and research use.