Enterprise Document Labeling & Classification (iFLYTEK)
Led the labeling and semantic classification of large-scale enterprise document datasets using AI-powered platforms. Developed and maintained prompt templates and structured output constraints for document knowledge extraction workflows. Built and optimized intelligent document parsing pipelines augmented with OCR for improved data accuracy. • Labeled over 100K documents daily for relevance, classification, and semantic similarity. • Designed enterprise RAG system pipelines to structure and process unstructured textual and scanned document data. • Integrated prompt engineering strategies for supervised fine-tuning of LLM responses using labeled data. • Employed Elasticsearch and custom internal tooling for document retrieval labeling and analytics.