Data Engineer Intern — Nube9 LLC (Feb 2024 – Apr 2024)
Engineered distributed ETL pipelines to prepare datasets for downstream analytics, involving data transformation and processing. Applied Apache Spark and Airflow to ingest and process 5TB+ of raw data weekly with a focus on improving data reliability. Optimized PostgreSQL SQL queries and indexing to increase execution speed by 40%, supporting higher-quality analytics inputs. • Built fault-tolerant ETL workflows and reduced processing time by 30%. • Eliminated 95% of pipeline failures through workflow redesign. • Coordinated with data scientists to refine data models for more accurate business reporting. • Used distributed processing patterns suitable for large-scale AI/analytics dataset preparation.