Azure Data Engineer · Mumbai, India

I build pipelines that turn raw data into something you can trust.

Two years designing and running Azure Data Factory and Databricks pipelines that move data from messy sources to analytics-ready gold tables — incrementally, with governance, and without surprises.

// how data actually moves through my pipelines
Source APIs, files, DBs
Bronze raw, incremental
Silver cleaned, typed
Gold star schema
Serving Synapse · Power BI
01

About

I'm an Azure Data Engineer working across data integration, pipeline design, and reporting for pharmacy and retail clients. Most of my day is spent in Azure Data Factory and Databricks — connecting sources, writing PySpark transformations, and making sure the same pipeline that runs today still runs cleanly a year from now.

I care about the parts that don't show up in a demo: incremental loads that don't reprocess everything, dimension tables that preserve history correctly, and pipelines built once and reused across datasets instead of copy-pasted five times.

Experience
2 years, Azure Data Engineering
Current role
Oberon Software Solution Pvt Ltd.
Focus
ETL/ELT, Delta Lake, Medallion Architecture
Based in
Mumbai, India
02

Skills

Cloud & Platforms
Azure Data FactoryAzure Databricks ADLS Gen2Azure Synapse Analytics Azure Key VaultAzure Logic Apps Azure Monitor
Data Engineering
Delta LakeDelta Live Tables Medallion ArchitectureStar Schema SCD Type 1 & 2CDC Data WarehousingMicrosoft Fabric
Programming
SQLPySparkPython
Tools
GitCI/CDDocker Power BIJupyter NotebookVS Code
03

Experience

Azure Data Engineer

Oberon Software Solution Pvt Ltd. — Pharmacy & Retail
Jul 2024 — Present
  • Gather requirements directly with clients, run data analysis, and contribute to solution design for new data engineering workloads.
  • Design and build ETL/ELT ingestion pipelines in Azure Data Factory and Azure Databricks, following the org's GDI framework.
  • Create ADF linked services to connect and standardize multiple source and target systems.
  • Write Spark SQL and PySpark transformations for complex aggregations and analytics workloads.
  • Built Structured Streaming ingestion with Databricks Auto Loader to process near-real-time Avro files.
  • Tuned Spark jobs with partitioning, caching, broadcast joins, and memory tuning.
30–50%reduction in Spark job processing time
04

Portfolio Projects

End-to-End Netflix Data Engineering Pipeline on Azure

BronzeSilverGold

A full ETL pipeline built on a public Netflix dataset — five source tables ingested through Azure Data Factory and Databricks, transformed with PySpark and Delta Lake, and modeled for reporting.

  • Dual ingestion: ADF pipeline (parameterized ForEach, source validation) lands dimension tables in Bronze; Databricks Auto Loader streams the larger, incrementally-growing fact table.
  • Medallion architecture (Bronze → Silver → Gold) with reusable, parameterized notebooks driven by Databricks Workflows instead of one notebook per dataset.
  • Integrated Power BI for real-time order trend visualization on top of the gold layer.
Azure Data FactoryDatabricks PySparkDelta Lake Power BI
</> View on GitHub →

Azure Databricks Retail Data Lakehouse

BronzeSilverGold

An end-to-end retail lakehouse across Orders, Customers, Products, and Regions — built to demonstrate governed, scalable dimensional modeling, not just a one-off script.

  • Delta Live Tables for declarative, quality-checked transformation across the medallion layers.
  • SCD-based dimension tables and star schema — surrogate keys and history preserved correctly across incremental loads, not just overwritten.
  • Automated end-to-end execution with Databricks Workflows and near-real-time ingestion via Spark Structured Streaming.
Azure DatabricksPySpark Delta LakeDelta Live Tables Auto Loader
</> View on GitHub →
05

Education

Bachelor of Engineering — Computer Engineering

University of Mumbai · 2024
CGPA 8.3 / 10