Data Labeling and Evaluation Specialist (AI Knowledge Base)
Developed document parsing and retrieval pipeline with vectorization and context assembly for intelligent Q&A AI. Labeled and curated training data for function calling tasks, creating question-context-response annotations for LLM evaluation. Conducted hands-on evaluation/ratings of LLM-generated answers with human feedback to improve performance. • Parsed and segmented documents into context chunks for embedding and retrieval testing. • Built function calling datasets for downstream internal API/tool invocation by LLMs. • Labeled ground truth responses and evaluated model output across multiple scenarios. • Performed ongoing annotation and model assessment for knowledge base Q&A and LLM function calls.