Staff Software Engineer
As a Staff Software Engineer, I designed and executed the architecture for large-scale distributed systems, focusing heavily on real-time streaming data ingestion and automated machine learning training pipelines. I engineered high-throughput, end-to-end data lakehouse solutions utilizing Apache Kafka for streaming ingestion and Apache Iceberg alongside Snowflake for scalable data storage. Leveraging C++, Python, and Scala, I programmed advanced backend layers that included automated data annotation pipelines to generate high-fidelity datasets for machine learning classification models. I also developed automated PII discovery and data redaction engines to ensure strict compliance with GDPR and CCPA privacy frameworks. To maintain rigorous quality standards, I enforced Test-Driven Development practices, utilized Semaphore CI for continuous integration, and implemented CQRS design patterns for optimal data consistency across the entire infrastructure.