Independent NLP & Annotation Research Project
You are designing a Swahili–English parallel annotation dataset project to support low-resource NLP model training, including researching annotation schemas such as IOB tagging and intent classification. This is direct work toward creating labeled datasets for AI training and evaluation, with emphasis on writing methodology and annotation guidelines. The project also includes self-directed development of data processing and preprocessing pipelines used in annotation workflows. • Build a parallel annotation dataset for Swahili–English NLP training. • Research annotation tools and schemas including IOB tagging, sentiment, and intent labels. • Develop supporting Python workflows using pandas and NLTK. • Document annotation guidelines and methodology using best practices for labeling.