Research Intern (Remote) — Data extraction and validation for LLM fine-tuning
Built an automated pipeline to extract structured information from 700+ MSHA Fatality Reports to support downstream model improvement. Validated the produced datasets using root-cause analysis to ensure the reliability and quality of the structured outputs for LLM fine-tuning. Organized and processed extracted fields into a consistent format for high-quality training data.• Source MSHA fatality reports were processed to derive structured records for model training.• Output validation used root-cause analysis to reduce errors and improve dataset trustworthiness.• Performance improvements reduced per-PDF handling time from ~4 minutes to ~0.5 seconds.• Created a dataset suitable for LLM fine-tuning workflows based on extracted structured content.