AI Training/Data Labeling – ControlLLM Corpus and Fine-Tuning
I was a core contributor to the design of the corpus and course-learning modules for a lightweight, general-purpose LLM (ControlLLM). This involved selecting, curating, and annotating large-scale text data to refine language model capabilities. The process centered on text data selection, prompt engineering, and supervised fine-tuning for LLM improvement.• Designed and curated LLM training corpora • Conducted text annotation and prompt engineering • Led fine-tuning of ControlLLM's language models • Focused on supervised learning strategies for model enhancement