Summary:
Overview
Work History
Skills
Certification
Timeline

Kanchan Diwate

Publicis Sapient
Pune
1
Certification
6
years of professional experience

Azure Data Engineer with 6.1 years of experience in designing and developing scalable cloud-based data engineering solutions on Microsoft Azure. Experienced in building end-to-end ETL/ELT pipelines using Azure Data Factory (ADF), Azure Databricks, PySpark, Spark SQL, Delta Lake, Azure Synapse Analytics, ADLS Gen2, Azure SQL Database, and SQL Server.

Hands-on experience with Lakehouse and Medallion Architecture (Bronze, Silver, Gold), Databricks Notebooks, Change Data Capture (CDC), incremental data loading, Delta Lake optimization, performance tuning, and Spark optimization techniques including partitioning, caching, broadcast joins, and AQE. Proficient in processing CSV, JSON, Parquet, REST APIs, and streaming data using Auto Loader and Structured Streaming.

Skilled in Python, PySpark, SQL, data modeling, Star & Snowflake Schema, data warehousing, and developing scalable, high-performance data pipelines. Experienced with Azure DevOps, Git, CI/CD, Unity Catalog, Delta Live Tables (DLT), Databricks Workflows, and Agile methodologies to deliver secure, reliable, and production-ready data solutions that support analytics and business intelligence.

Work History

Senior Associate Data Engineer L2

10 Months
Publicis Sapient | 05.2025 - 03.2026

Roles & Responsibilities

Project 1: ASO Hardline Imagery – India (Dec/2025-Mar/2026)

Summary: This project was focused on creating high-quality, realistic product imagery using AI-based image generation workflows. Worked extensively with master prompts and negative prompts to generate accurate visuals aligned with universal guidelines and product-specific client requirements. Real product images from the client website were analyzed to design custom prompts, ensuring that the generated images matched brand expectations and met quality standards. I also performed detailed quality testing by comparing system-generated and manually generated images with actual stock images and documented all status updates through JIRA for seamless project tracking. Prompt Engineering (Master + Negative Prompts)

AI Models / Image Generation Tools: Vortex AI, Nano Banana, Gemini Flash, Nano Banana Pro

Platforms / Input Sources: Client Website Product Images

Image Quality Standards: Universal guidelines + product-specific guidelines Project Tracking.

JIRA Responsibilities:

• Developed custom master and negative prompts to generate high-quality, realistic product images aligned with client guidelines and brand expectations.

• Analyzed client website product images and created tailored prompts to match color, texture, design, perspective, and product detailing.

• Generated product visuals using Vortex AI, Nano Banana, Gemini Flash, and Nano Banana Pro based on universal and product-specific guidelines.

• Performed quality testing and validation of system-generated and manually created images by comparing them with actual stock images.

• Ensured all images met realism, accuracy, clarity, lighting, and consistency standards before submission to the client.

• Maintained precise and timely JIRA updates for image generation tasks, testing outcomes, and approval workflows.

• Image Quality Evaluation & Validation

• AI-based Realistic Image Generation Manual vs System Image Comparison

Project 2: Health care domain project (May /2025- Dec 2025)

Worked as an Azure Data Engineer on a healthcare data platform modernization project focused on building scalable, secure, and high-performance data pipelines for enterprise analytics. Designed and implemented end-to-end Azure data engineering solutions using Azure Data Factory, Azure Databricks, Delta Lake, and Azure Data Lake Storage Gen2. Built curated Bronze, Silver, and Gold data layers, enabling trusted datasets for reporting, advanced analytics, and business intelligence.

Environment: Azure Data Factory (ADF), Azure Databricks, PySpark, Spark SQL, Delta Lake, Azure Data Lake Storage Gen2 (ADLS Gen2), Azure SQL Database, Azure Synapse Analytics, Azure DevOps, Git, Unity Catalog, Databricks Workflows, JIRA

  • Developed scalable ETL/ELT pipelines using Azure Data Factory (ADF) and Azure Databricks for healthcare data integration.
  • Designed and implemented Bronze, Silver, and Gold layers using Delta Lake following the Medallion Architecture.
  • Ingested healthcare data from Azure SQL Database, SQL Server, CSV, JSON, REST APIs, and Parquet into ADLS Gen2.
  • Built transformation logic using PySpark and Spark SQL for cleansing, validation, standardization, enrichment, and business rule implementation.
  • Implemented CDC, incremental loading, schema evolution, and Slowly Changing Dimension (SCD) processing.
  • Optimized large-scale Spark jobs using partitioning, caching, broadcast joins, AQE, Delta OPTIMIZE, Z-ORDER, and VACUUM to improve performance and reduce execution time.
  • Created reusable Databricks notebooks and orchestrated end-to-end workflows using ADF pipelines and Databricks Workflows.
  • Implemented comprehensive data quality checks, reconciliation processes, and exception handling to ensure reliable data delivery.
  • Managed code repositories using Git, automated deployments through Azure DevOps CI/CD, and maintained environment-specific configurations.
  • Configured Unity Catalog for data governance, role-based access control, and secure data sharing.
  • Collaborated with business analysts, data architects, and reporting teams to deliver production-ready Gold datasets for downstream analytics and Power BI reporting.
  • Participated in Agile ceremonies, production support, release deployments, and issue tracking using JIRA.

Programmer Analyst

4 Months
BITWISE | 11.2024 - 03.2025

Project: Enterprise Data Platform Modernization Summary Tools & Technologies Responsibilities

Worked as an Azure Data Engineer on an enterprise-wide cloud data platform modernization initiative to build scalable, secure, and high-performance data pipelines for analytics and reporting. Designed and implemented modern Lakehouse Architecture on Microsoft Azure using Azure Data Factory (ADF), Azure Databricks, Delta Lake, Azure Synapse Analytics, and ADLS Gen2. Developed end-to-end data ingestion, transformation, orchestration, monitoring, and governance solutions to deliver trusted, analytics-ready datasets while improving pipeline reliability, scalability, and operational efficiency.

Azure Data Factory (ADF), Azure Databricks, PySpark, Spark SQL, Delta Lake, Azure Synapse Analytics, Azure Data Lake Storage Gen2 (ADLS Gen2), Azure SQL Database, Azure DevOps, Git, Unity Catalog, Databricks Workflows, Delta Live Tables (DLT), Auto Loader, Structured Streaming, REST APIs, SQL, Python, JIRA

  • Designed and implemented enterprise-scale Lakehouse Architecture on Azure to support high-volume batch and near real-time data processing.
  • Built metadata-driven ADF pipelines for orchestrating complex data ingestion workflows from enterprise applications, databases, cloud storage, and REST APIs.
  • Developed reusable PySpark frameworks and modular Databricks notebooks to standardize transformation logic across multiple business domains.
  • Implemented Delta Live Tables (DLT) to automate data quality enforcement, schema validation, expectation rules, and reliable pipeline execution.
  • Built scalable ingestion frameworks using Auto Loader with schema inference and schema evolution to efficiently process continuously arriving files.
  • Designed and optimized Delta Lake tables using appropriate partitioning strategies, file compaction, liquid clustering/Z-Ordering (where applicable), and lifecycle maintenance to improve query performance and storage efficiency.
  • Implemented enterprise data governance using Unity Catalog, including fine-grained access control, data lineage, catalog management, and secure data sharing.
  • Developed automated monitoring, logging, and alerting mechanisms for pipeline health, job failures, SLA tracking, and operational support.
  • Created reusable parameterized notebooks and configuration-driven frameworks to support multiple environments and reduce development effort.
  • Automated code deployment and release management using Azure DevOps, Git branching strategies, and CI/CD pipelines for development, testing, and production environments.
  • Collaborated with solution architects, data analysts, and downstream reporting teams to deliver certified, analytics-ready datasets for enterprise reporting and advanced analytics.
  • Participated in production support, root cause analysis, performance optimization, release planning, and Agile sprint activities while ensuring compliance with enterprise data engineering standards.

Azure Data Engineer

4 Years 3 Months
DATA DELVE TECHNOLOGIES PRIVATE LIMITED | 07.2020 - 10.2024

Project 1: Media Domain Data Engineering Platform (Mar 2023 – Oct 2024) Project 2: Insurance Data Integration Platform (Aug 2021 – Feb 2023)

Responsibilities:

  • Developed scalable ETL/ELT pipelines using Azure Data Factory (ADF) and Azure Databricks.
  • Built PySpark and Spark SQL notebooks for data ingestion and transformation.
  • Implemented incremental data loading using metadata-driven frameworks.
  • Integrated MySQL metadata with ADF Lookup activities for dynamic pipeline execution.
  • Optimized Spark jobs through partitioning, caching, and query tuning.
  • Automated workflow scheduling using ADF Triggers and Databricks Workflows.
  • Implemented logging, error handling, monitoring, and production support.
  • Managed source code and deployments using Git and Azure DevOps.

Responsibilities:

  • Designed cloud-based ETL pipelines using ADF, Azure Databricks, and PySpark.
  • Ingested and transformed policy, claims, and customer data from multiple sources.
  • Implemented data validation, cleansing, deduplication, and business rules.
  • Developed optimized Spark SQL and SQL queries for high-performance data processing.
  • Created metadata-driven and reusable pipeline components.
  • Performed data reconciliation and source-to-target validation.
  • Supported CI/CD deployments using Azure DevOps and Git.
  • Collaborated in Agile teams and provided production support for enterprise data pipelines.

Azure Data Engineer Trainee

3 Months
04.2021 - 07.2021

Designation: Responsibilities:

  • Assisted in developing ETL/ELT pipelines using Azure Data Factory (ADF) and Azure Databricks.
  • Gained hands-on experience in building PySpark and Spark SQL notebooks for data ingestion and transformation.
  • Loaded and processed data from CSV, JSON, SQL Server, and Azure Data Lake Storage Gen2 (ADLS Gen2).
  • Performed data cleansing, transformation, validation, and basic data quality checks using PySpark and SQL.
  • Assisted in implementing the Medallion Architecture (Bronze, Silver, Gold) for structured data processing.
  • Wrote SQL queries for data extraction, validation, reconciliation, and troubleshooting ETL workflows.
  • Supported pipeline scheduling, monitoring, and issue resolution using Azure Data Factory and Databricks Workflows.
  • Used Git and Azure DevOps for source code management and deployment support.
  • Collaborated with senior data engineers in Agile/Scrum teams to develop scalable and reliable cloud data solutions.

Skills

Data Modeling
Data Transformation
Data Analysis
Power BI
Data Visualization
Azure Databricks
Azure Data Factory
PySpark
Spark SQL
Delta Lake
Azure Synapse
ADLS Gen2
ETL/ELT
SQL
Azure DevOps -CI/CD
Unity Catalog
Delta Lake
Lakehouse Medallion Architecture
Data Warehousing

Certification

Microsoft Certified: Azure Data Engineer Associate_DP203

Timeline

Senior Associate Data Engineer L2

Publicis Sapient
05.2025 - 03.2026Read More

Programmer Analyst

BITWISE
11.2024 - 03.2025Read More

Azure Data Engineer Trainee

04.2021 - 07.2021Read More

Azure Data Engineer

DATA DELVE TECHNOLOGIES PRIVATE LIMITED
07.2020 - 10.2024Read More
Kanchan Diwate