Summary
Data Engineer with 5 years building scalable ETL pipelines and data platforms on AWS and GCP. Expert in Spark, Airflow, and Snowflake. Designed streaming infrastructure processing 4TB/day and cut pipeline costs by 35% at Expedia.
Experience
Architected a Spark streaming platform processing 4TB/day, reducing data latency from 6 hours to under 5 minutes.
Migrated 80+ batch jobs from on-prem Hadoop to Snowflake and dbt, cutting compute costs by 35% ($310K annually).
Built Airflow orchestration with automated data-quality checks, reducing pipeline failures by 62%.
Developed ETL pipelines in Python and Airflow ingesting 200M+ daily events into a Redshift warehouse.
Optimized partitioning and clustering strategies, improving analyst query performance by 4x.
Implemented a metadata catalog with data lineage, cutting onboarding time for new analysts by 40%.
Education
Projects
Built a Kafka and Spark Structured Streaming pipeline scoring 50K transactions/second, flagging fraud with 99.2% precision and saving an estimated $1.2M in chargebacks.
Authored a dbt macro library adopted by 1,400+ GitHub users, automating incremental model testing and snapshot validation.