
Lead Data Engineer with 10+ years of experience designing, developing, and optimizing scalable cloud-native data platforms, ETL/ELT pipelines, and distributed data processing solutions. Expertise in Databricks, Snowflake, PySpark, Apache Spark, Python, SQL, BigQuery, Kafka, Airflow, Hive, AWS, and GCP, delivering high-performance batch and real-time data pipelines for multi-terabyte datasets. Proven success in modernizing legacy data ecosystems, building enterprise data lakes and warehouses, and implementing cloud solutions that reduced infrastructure costs by 30% and improved pipeline performance by 40%. Strong experience in data architecture, business intelligence, cloud migration, data governance, CI/CD, DevOps, and performance optimization. Skilled at leading cross-functional engineering teams, mentoring developers, collaborating with stakeholders, and delivering secure, scalable, and reliable data solutions that drive business value.