Independent NLP & Annotation Research Project
Self-directed research to design a Swahili–English parallel annotation dataset intended for low-resource NLP model training. Studied industry-standard annotation tools and tagging schemas, including IOB tagging, sentiment labels, and intent classification, to inform label taxonomy and guideline creation. Documented annotation methodology and quality practices to support reliable dataset construction for model development. • Designing annotation dataset for cross-lingual NLP • Researching labeling tools and schemas (e.g., IOB, sentiment, intent) • Implementing NLP preprocessing workflows with Python • Writing annotation guidelines for consistent labeling