Two years designing and running Azure Data Factory and Databricks pipelines that move data from messy sources to analytics-ready gold tables — incrementally, with governance, and without surprises.
I'm an Azure Data Engineer working across data integration, pipeline design, and reporting for pharmacy and retail clients. Most of my day is spent in Azure Data Factory and Databricks — connecting sources, writing PySpark transformations, and making sure the same pipeline that runs today still runs cleanly a year from now.
I care about the parts that don't show up in a demo: incremental loads that don't reprocess everything, dimension tables that preserve history correctly, and pipelines built once and reused across datasets instead of copy-pasted five times.
A full ETL pipeline built on a public Netflix dataset — five source tables ingested through Azure Data Factory and Databricks, transformed with PySpark and Delta Lake, and modeled for reporting.
An end-to-end retail lakehouse across Orders, Customers, Products, and Regions — built to demonstrate governed, scalable dimensional modeling, not just a one-off script.