Senior Big Data Engineer
As a Senior Big Data Engineer, I lead the design and deployment of large-scale data pipelines, leveraging advanced technologies across cloud platforms. I am responsible for architecting and optimizing real-time analytics systems that transform massive amounts of data into meaningful business insights. I also collaborate with cross-functional teams to ensure the scalability and efficiency of our solutions, while mentoring junior engineers in industry best practices. • Designed and deployed data pipelines using Apache Spark and Python to process 80M+ events daily across AWS and GCP. • Architected a real-time streaming platform with Kafka and Spark Structured Streaming, reducing reporting latency from 20 minutes to under 8 seconds. • Enabled productionization of ML models in PySpark for automated customer segmentation on 15M+ profiles. • Optimized Hive and BigQuery models, cutting compute costs by 45% and query runtime by 38%.