For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
P
Pranay B.

Pranay B.

Research Intern (Remote) — Data extraction and validation for LLM fine-tuning

India flagWardha, India

Key Skills

Software

No software listed

Top Subject Matter

Mine safety incident reporting (MSHA Fatality Reports)
Mine safety root-cause analysis and regulatory compliance (DGMS/MSHA)
Legal Services & Contract Review

Top Data Types

DocumentDocument
VideoVideo
TextText

Top Task Types

Fine-tuningFine-tuning

Freelancer Overview

Research Intern (Remote) — Data extraction and validation for LLM fine-tuning. Brings 1+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include PyMuPDF, Pandas, and Unsloth. Education includes Dual Bachelor of Technology and Master of Technology, Indian Institute of Technology Kharagpur (2026). AI-training focus includes data types such as Document and labeling workflows including Evaluation, Rating, and Fine-tuning.

Labeling Experience

Research Intern (Remote) — Data extraction and validation for LLM fine-tuning

DocumentDocument

Built an automated pipeline to extract structured information from 700+ MSHA Fatality Reports to support downstream model improvement. Validated the produced datasets using root-cause analysis to ensure the reliability and quality of the structured outputs for LLM fine-tuning. Organized and processed extracted fields into a consistent format for high-quality training data.• Source MSHA fatality reports were processed to derive structured records for model training.• Output validation used root-cause analysis to reduce errors and improve dataset trustworthiness.• Performance improvements reduced per-PDF handling time from ~4 minutes to ~0.5 seconds.• Created a dataset suitable for LLM fine-tuning workflows based on extracted structured content.

2025 - 2025

Master Thesis Project — LLM fine-tuning + RAG for RCA generation

DocumentDocumentFine-tuningFine-tuning

Fine-tuned a Llama 3.2 3B-Instruct model to generate structured root-cause analysis (RCA) reports from mine incident narratives for DGMS compliance. Used Unsloth with QLoRA on 700+ MSHA reports to adapt the model for domain-specific instruction following and structured output generation. Integrated a RAG pipeline with citation-grounded answers using LangChain, Nomic embeddings, and ChromaDB to support regulatory citations in responses.• Training used QLoRA fine-tuning on a dataset of 700+ MSHA reports.• Objective was to convert narrative incidents into structured RCA reports with regulatory citations.• RAG components enabled citation-grounded answers using LangChain with Nomic embeddings and ChromaDB.• Built an interactive Gradio interface to produce structured RCA reports quickly (under 5 seconds).

2026

Education

I

Indian Institute of Technology Kharagpur

Dual Bachelor of Technology and Master of Technology, Mining and Safety Engineering

Dual Bachelor of Technology and Master of Technology
2021 - 2026

Work History

U

University of Utah

Research Intern

Salt Lake City
2025 - 2025