Intern - PointCross Life Sciences (ETL pipelines for PDF document data extraction)
Developed Python-based ETL pipelines to extract, clean, and transform structured data from unstructured PDF documents according to client-specific business requirements. Automated document processing workflows using PDF parsing to reduce manual data extraction effort and improve operational efficiency. Conducted data cleaning, preprocessing, validation, and transformation using Pandas and NumPy to prepare datasets for downstream analytical workflows. • Extracted and transformed structured data from unstructured PDFs. • Automated PDF parsing-based document processing. • Performed cleaning, validation, and transformation with Pandas/NumPy. • Built reusable configurable extraction scripts for multiple document formats.