Transforming raw data into strategic insights. Building scalable pipelines that drive revenue and power decisions across retail, e-commerce, and pharma.
Data Engineer with 5+ years of experience in analytics and pipeline development across retail, e-commerce, and pharmaceutical industries.
I build scalable ELT pipelines using Python, PySpark, and dbt on AWS and Snowflake, delivering data models that drive revenue, improve accuracy, and reduce defects. My work has directly contributed to recovering $1M+ GMV and cutting manual reporting by 90%.
Data Analytics & Information Systems
Open to remote & relocation
class DataEngineer:
def __init__(self):
self.name = "Baanu Sai"
self.role = "Data Engineer"
self.stack = [
"Python", "PySpark",
"Snowflake", "dbt",
"AWS", "SQL"
]
def build_pipeline(self):
return "Scalable & Reliable"
Modern e-commerce data warehouse with Medallion Architecture (Bronze/Silver/Gold), dimensional modeling, incremental processing, dbt tests, and DuckDB-powered local development.
End-to-end real-time pipeline ingesting events through Kinesis → Lambda → Firehose into S3, with Snowpipe auto-ingest and Snowflake Streams & Tasks driving incremental downstream transforms.
Reusable PySpark + Great Expectations data quality framework with YAML-driven expectations, pluggable reporters, and CI-friendly validation runs.
Fully serverless ETL with Lambda ingestion, Step Functions orchestration, Glue transformations, DynamoDB state tracking, and Terraform IaC deployment.
Complete data lakehouse with Delta Lake ACID transactions, time travel, schema evolution, MERGE upserts, and Z-Order optimization.
Amazon Web Services
LinkedIn Learning
Looking for a Data Engineer who can turn complex data challenges into elegant solutions? Let's connect.