LLM data curation & supervised fine-tuning data preparation (security-focused training data)
Built and prepared LLM training inputs for security-oriented applications by crafting prompt-response style data and curated examples for model behavior. Focused on generating and structuring high-quality training materials to improve accuracy and safety in threat-detection contexts. Worked on data curation practices including filtering, deduplication, and synthetic data generation for supervised workflows. • Created prompt-response pairs tailored to cybersecurity use cases. • Performed knowledge/data filtering and deduplication steps to improve training quality. • Generated synthetic data to expand coverage for LLM training. • Prepared safety/accuracy-oriented evaluation-ready examples.