Project West-Voice: Multi-Dialect Nigerian English Corpus for Low-Resource ASR
Project Scope The scope of Project West-Voice was to establish a high-quality, linguistically diverse, and time-aligned acoustic-textual dataset specifically for multi-dialect Nigerian English and English/Pidgin code-switching. The project aimed to address a significant bottleneck in speech technology: the high error rates of standard global Automatic Speech Recognition (ASR) models when processing West African phonetic variations, stress patterns, and localized vocabulary. The scope encompassed the ingestion of raw field audio, audio pre-processing and splitting, multi-layer semantic and acoustic annotation, and packaging into a machine-learning-ready corpus optimized for lightweight, low-resource automatic speech recognition architectures. Data Labelling Tasks Performed The data annotation pipeline was structured into three distinct layers to ensure multi-dimensional training utility: Time-Aligned Verbatim Transcription: Annotators performed precision audio-to-text mapping, transcribing conversational fragments word-for-word. This included capturing speech fillers, repetitions, and localized phonemes, linked directly to exact structural time-markers (timestamps) to align speech frequencies with text tokens. Acoustic & Environmental Metadata Tagging: Every audio segment was evaluated and tagged with categorical attributes. Labelers classified primary linguistic influences (e.g., L1-Yoruba, L1-Hausa, L1-Igbo, or Pidgin-dominant), speaker demographics (gender and approximate age bracket), and background acoustic environments (e.g., clean studio, low ambient street noise, high echo). Linguistic Named Entity Recognition (NER): High-level linguistic mapping was applied to isolate code-switched phrases and distinct West African English loanwords or slang. These localized expressions were systematically tagged and mapped into a structured JSON schema to train downstream language models on contextual syntax. Project Size Dataset Volume: A robust corpus comprising 500+ hours of localized, high-fidelity audio data (.wav and .flac formats sampled uniformly at 16kHz/mono-channel) along with corresponding structural JSON text payloads. Demographic & Regional Representation: Captured distinct speech profiles from over 1,200 unique speakers across diverse urban and rural hubs in Nigeria, balancing regional accents to ensure comprehensive acoustic coverage. Annotation Points: The dataset comprised millions of individual time-aligned word tokens, hundreds of thousands of segmented audio boundaries, and exhaustive multi-attribute metadata logs mapping linguistic variations. Measures Adhered To Acoustic Quality Baseline: Adhered to strict signal-to-noise ratio (SNR) thresholds during data filtering. Raw audio frames that fell below strict quality standards or contained completely overlapping unintelligible voices were automatically culled or flagged for re-recording to ensure clean model features. Data Privacy & PII Compliance: In alignment with data protection regulations such as the Nigeria Data Protection Act (NDPA) and GDPR frameworks, all audio segments were thoroughly audited. Any accidentally captured Personally Identifiable Information (PII)—including specific addresses, banking mentions, or full legal names—was meticulously redacted or masked before being sent to the annotation pipeline. Inter-Annotator Agreement (IAA) & Quality Assurance: To maintain rigorous labeling accuracy, a multi-stage validation matrix was enforced. Labeling files were subjected to blind cross-verification by secondary annotators, maintaining a target Inter-Annotator Agreement (IAA) score of >95% (via Cohen's Kappa metric) before passing into the final deployment-ready training corpus.