LLM Distiller (Data Curation & Alignment) — open-source distillation and alignment project.
Developed an open-source LLM distillation workflow focused on high-fidelity dataset curation and fine-tuning alignment to preserve benchmark performance. Distilled large 70B-parameter models down to 7B parameters while maintaining evaluation quality within a small performance margin. Curated and aligned training data to support downstream model behavior targets. • Curated datasets for high-fidelity distillation. • Performed fine-tuning alignment to preserve benchmark performance. • Built an LLM distiller pipeline for 70B→7B compression. • Validated performance retention within an estimated 3% margin.