Text Data Annotation and LLM Training for AI Knowledge Assistant
Curated, chunked, and annotated domain-specific text data for AI knowledge retrieval tasks supporting an LLM-based assistant. Extracted metadata and semantic features from technical documents, textbooks, and past exam papers to train and evaluate retrieval-augmented models. Assessed and classified text segments for relevance, knowledge coverage, and context coherence in support of question-answering and knowledge assistance features. • Labeled and extracted meaningful sections from computer science educational texts. • Implemented semantic similarity retrieval and metadata annotation strategies. • Supported LLM fine-tuning and evaluation with high-quality text data annotation. • Used PGvector and proprietary databases for text data labeling and management.