LLM Prompt & AI Agent Training Data Annotation Project
This project builds high-quality training datasets for large language models and customized AI agents. My core tasks include writing, screening, scoring and optimizing system prompt samples, establishing unified annotation specifications for multi-turn dialogue data, auditing data logic accuracy and eliminating invalid low-quality entries. The dataset scale covers tens of thousands of prompt and dialogue samples. I adopted double-check quality control standards to ensure data uniformity, which greatly boosted model fine-tuning effect and stable operation of deployed intelligent agents. I also adjusted annotation rules iteratively according to model test feedback, and connected labeled datasets with full-stack development and online deployment workflows.Main labeling work covers text generation, prompt crafting, coding dataset sorting and RLHF evaluation rating for LLM fine-tuning.