AI Data Annotation and LLM Fine-tuning (Ongoing Project)
In the ongoing AI sales and video generation project, I participated in preparing labeled data for fine-tuning deepseek-32b models. Data sources included sales Q&A scripts and viral video scripts, where initial data extraction was followed by human annotation and AI-assisted filtering. The objective was to produce a high-quality Alpaca-compatible dataset for improved local LLM fine-tuning using internal pipelines. • Utilized Dify for sales dialogue extraction and initial data curation. • Performed manual annotation for quality control and consistency of the dataset. • Collaborated in leveraging AI methods for semi-automated data cleansing and augmentation. • Final dataset was used to fine-tune the deepseek-32b distilled model using llama-factory.