Synthetic Data Generation & Model Distillation - Instruction Data Curation
I designed synthetic data generation pipelines leveraging GPT-4 and other open-source LLMs to produce instruction samples for legal and financial Q&A. I applied semantic filtering to refine generated data, ensuring high similarity and quality for downstream training. My work enabled efficient model distillation and boosted student model effectiveness. • Generated and filtered 50,000+ synthetic instruction data points • Automated sample selection using FAISS-based semantic similarity for consistency • Created datasets that substantially improved student-teacher knowledge transfer • Deployed and maintained data generation systems in Dockerized environments