Shubham Yadav
Mumbai, IndiaAvailable for opportunities

Senior Data EngineerSAS · PySpark · Azure Databricks · GCP · Snowflake

Building modern data platforms from complex data.

I'm a Data Engineer Professional with 7+ years of experience in data engineering, SAS, ETL, data warehousing and analytics. My work spans traditional enterprise data platforms and modern cloud-native technologies with a focus on scalable pipelines, legacy modernization and reliable data platforms.

SASPySparkApache SparkDatabricksDelta LakeGCPSnowflakeSQL
SAS / Legacy01ETL / PySpark02Lakehouse03Snowflake04

7+ Years

Data Engineering Experience

SAS

Enterprise ETL & Data Warehousing

Cloud

GCP & Azure

Modern Data

PySpark · Databricks · Delta Lake

Data Platforms

Snowflake · PostgreSQL · Oracle · MySQL

Visualization & Reporting

SAS Viya 3.5

About

Turning enterprise data into modern data platforms.

I started my career working extensively with SAS, ETL, data warehousing and enterprise reporting. Over the years, I've expanded into modern data engineering and cloud technologies PySpark, Azure Databricks, Delta Lake, GCP and Snowflake.

I enjoy working on complex data problems from migrating legacy platforms and designing ETL pipelines to building scalable cloud-based data processing solutions.

My approach combines strong fundamentals in data engineering with modern cloud technologies, bridging traditional enterprise systems and next-generation data platforms.

What I specialize in

  • Data Engineering
  • ETL / ELT Development
  • Cloud Data Migration
  • Data Platform Modernization
  • Distributed Data Processing
  • Data Warehousing
  • Lakehouse Architecture
  • Data Governance
  • Data Quality
  • Enterprise Analytics

Technical Expertise

A stack built across two eras of data engineering.

Enterprise SAS foundations, extended into modern distributed and cloud-native platforms.

Data Engineering

  • Python
  • PySpark
  • Apache Spark
  • SQL
  • ETL / ELT
  • Data Warehousing
  • Data Transformation
  • Data Quality
  • Distributed Processing
  • CI/CD

SAS

  • Base SAS
  • Advanced SAS
  • SAS DI Studio
  • SAS Viya 3.5
  • SAS CI360
  • SAS Intelligent Decisioning

Databricks

  • Azure Databricks
  • PySpark
  • Delta Lake
  • Delta Live Tables / Lakeflow
  • Auto Loader
  • Unity Catalog
  • Delta Sharing
  • Medallion Architecture
  • SCD Type 1 / Type 2
  • Data Governance

Google Cloud

  • Google Cloud Platform
  • Cloud Storage / GCS
  • Dataproc
  • Cloud SQL PostgreSQL
  • Storage Transfer Service
  • GCP Data Migration
  • Spark on Dataproc

Azure

  • Azure Databricks
  • ADLS Gen2
  • Azure SQL
  • Azure Data Engineering

Snowflake

  • Snowflake
  • Snowpipe
  • Streams
  • Tasks
  • Stored Procedures
  • Data Sharing
  • Medallion Architecture

Databases

  • PostgreSQL
  • Oracle
  • MySQL
  • Azure SQL
  • Snowflake

Governance & Security

  • Unity Catalog
  • RBAC
  • Row-Level Security
  • Column-Level Security
  • Data Sharing
  • Audit & Monitoring
  • Data Quality

Featured Experience

Senior Consultant / Senior Data Engineer

Working across enterprise data engineering, SAS modernization and cloud data platforms.

Focus Area

Enterprise Data Engineering & Cloud Modernization

  • Developing and maintaining enterprise ETL pipelines.
  • Working with SAS and SAS DI Studio for data integration.
  • Design the Data Model for Silver and Gold Layer(Dimensional Modeling)
  • Designing data transformation workflows using PySpark.
  • Working with Databricks and Delta Lake.
  • Building and optimizing cloud-based data processing solutions.
  • Supporting migration of legacy data workloads to cloud platforms.
  • Working with GCP services including Dataproc, GCS and Cloud SQL.
  • Implementing data validation and quality checks.
  • Troubleshooting performance, schema and data processing issues.
  • Working with modern data architecture and governance practices.

Featured Projects

Migrations and platforms, end to end.

Three representative builds — from legacy SAS modernization to lakehouse architecture and Snowflake automation.

Project 01

SAS Viya → GCP Modernization

Modernizing enterprise SAS workloads for the cloud.

A large-scale modernization initiative focused on migrating SAS data and workloads on a GCP Platform.

SAS ViyaStorage Transfer Service/Data FlowGoogle Cloud StorageDataproc / PySparkCloud SQL PostgreSQL

Technology Stack

SAS ViyaParquetGoogle Cloud StorageDataprocPySparkCloud SQL PostgreSQLStorage Transfer ServiceCLoud ComposerData Flow

Project Focus

Modernize legacy data workflows while maintaining data integrity, reliability and compatibility with downstream processes.

Key Areas

  • Source System Understanding
  • Gathering the Data and Job Inventory
  • SAS dataset modernization
  • Parquet conversion
  • Storage Transfer Service
  • Cloud data migration
  • GCS-based storage
  • PySpark transformations
  • Data validation
  • PostgreSQL-based curated data layer

Project 02

Databricks Lakehouse Platform

Building scalable data pipelines using modern lakehouse architecture.

Designed and worked with a Medallion Architecture consisting of Bronze, Silver and Gold layers.

SourcesAutoloader/API/LakeflowBronzeSilverGold(Facts & Dimension Tables)Power BI / Data Sharing

Technology Stack

DatabricksPySparkDelta LakeDLT / LakeflowAuto LoaderUnity CatalogDelta SharingCI/CD

Key Areas

  • Solution Design
  • Silver / Gold Data Modeling
  • Technical Implementation
  • Incremental data ingestion
  • Auto Loader
  • Streaming and batch processing
  • Schema evolution
  • Delta Lake
  • SCD Type 1 / Type 2
  • Data quality
  • Unity Catalog
  • Data governance
  • Data sharing
  • Leading the team of size 3 developer

Project 03

Snowflake Data Engineering POC

Exploring modern Snowflake data engineering patterns.

Designed a proof-of-concept covering incremental ingestion, change processing and workflow automation.

Source DataSnowpipeStreamsTasks / Stored ProceduresCurated Tables

Technology Stack

SnowflakeSnowpipeStreamsTasksStored Procedures

Key Areas

  • Continuous data ingestion
  • Incremental processing
  • Change data capture patterns
  • Automated workflows
  • Data transformation
  • Medallion-style data modeling

Project 04

Data Warehousing & Reporting

Centralized Data storage solution for analysis, reporting and machine learning use case

Designed and maintained enterprise ETL pipelines using SAS DI and SAS Viya for data warehousing and reporting.

Source DataSAS DIRawStageWarehouseMart/Reporting/Analysis

Technology Stack

SAS programming SAS DISAS MarcosOracle ProceduresSAS Viya 3.5Dashboard Reporting

Key Areas

  • ETL pipeline Development
  • Data Validation
  • Oracle to SAS Conversion
  • Job Optimization
  • Business Management Dashboard Creation
  • API Creationg for Data Tranfer

Domain Experience

Regulated, data-heavy industries.

Insurance

Experience working with enterprise insurance data, analytics and modernization initiatives.

Banking

Experience working with enterprise banking data and data processing workflows.

Education & Certifications

Formal grounding, verified skills.

Education

IIT Jodhpur

Convocation — June 2026

Postgraduate Diploma in Data Engineering

Distributed SystemsData EngineeringCloud TechnologiesBig Data ProcessingModern Data Platforms

University of Mumbai

Completed — 2019

Bachelor's Degree in Information Technology

Certifications

SnowPro Core

Snowflake

Snowflake Data Platform and Data Engineering

View certificate
SnowPro Core certificate

DP-203

Microsoft Azure

Azure Data Engineer Associate

Contact

Let's build something meaningful with data.

If you're working on data engineering, cloud modernization or modern data platforms, I'd be happy to connect.