Data Engineer

Baanu Sai Sankar Bojja

Transforming raw data into strategic insights. Building scalable pipelines that drive revenue and power decisions across retail, e-commerce, and pharma.

0 + Years Experience
0 M+ GMV Recovered
0 % Reporting Time Cut
Scroll to explore

Crafting Data Solutions
That Matter

Data Engineer with 5+ years of experience in analytics and pipeline development across retail, e-commerce, and pharmaceutical industries.

I build scalable ELT pipelines using Python, PySpark, and dbt on AWS and Snowflake, delivering data models that drive revenue, improve accuracy, and reduce defects. My work has directly contributed to recovering $1M+ GMV and cutting manual reporting by 90%.

🎓

Texas State University

Data Analytics & Information Systems

GPA: 3.7 | 2024
📍

Based in Texas

Open to remote & relocation

pipeline.py
class DataEngineer:
    def __init__(self):
        self.name = "Baanu Sai"
        self.role = "Data Engineer"
        self.stack = [
            "Python", "PySpark",
            "Snowflake", "dbt",
            "AWS", "SQL"
        ]
    
    def build_pipeline(self):
        return "Scalable & Reliable"

Skills & Technologies

Languages & Frameworks

SQL Python PySpark Spark SQL Scala JavaScript dbt

Data Platforms

Snowflake AWS Redshift AWS Athena BigQuery Kafka

Cloud & Infrastructure

AWS Glue AWS Lambda AWS S3 Terraform Docker Airflow dbt

Visualization & BI

Power BI Tableau QuickSight DAX

Concepts

ETL/ELT Data Modeling Star Schema CDC Data Quality A/B Testing

DevOps & Tools

Git REST APIs CI/CD CloudWatch

Professional Experience

Data Engineer

Real Value Products – Value RX
San Antonio, TX June 2024 – Present
  • Built end-to-end ELT pipelines in Python & PySpark with Bronze/Silver/Gold medallion architecture — improving data accuracy by ~20%
  • Modeled retail analytics data warehouse using dbt on Snowflake with star/snowflake schemas
  • Optimized SQL Server stored procedures — reduced query times from minutes to <10 seconds
  • Implemented automated data-quality checks — reduced listing defects by ~30%
  • Designed Power BI executive dashboards with RLS and custom DAX measures
PythonPySparkdbtSnowflakePower BI

Data Engineering Assistant

Texas State University
San Marcos, TX June 2022 – May 2024
  • Built reusable SQL/Python pipelines — cut weekly report prep time by ~60%
  • Developed interactive Tableau/Power BI dashboards for enrollment and retention tracking
  • Conducted statistical analyses (A/B tests, regression) to evaluate tutoring programs
  • Implemented data-quality checks and created lightweight data dictionary
SQLPythonTableauPower BI

Business Intelligence Analyst

Amazon
Bengaluru, India March 2020 – August 2022
  • Built automated vendor health dashboard — reduced manual reporting by 90% and cut SLA breaches by 28%
  • Created Python ETL pipelines normalizing SP-API feeds across 200+ vendors — increased anomaly detection precision by 35%
  • Designed SQL models tracking Buy Box inhibitors — recovered ~$1.2M GMV
  • Led root-cause analyses reducing chargeback incidence by 22%
  • Built alerting pipeline — cut time-to-detection from 12h to <30m, preventing ~$400k in margin leakage annually
PythonAWS GlueLambdaRedshiftAthenaQuickSight

Featured Projects

Certifications

AWS Certified Solutions Architect

Amazon Web Services

Data Analytics Professional

Google

NoSQL Essential Training

LinkedIn Learning

Let's Build
Something Great

Looking for a Data Engineer who can turn complex data challenges into elegant solutions? Let's connect.

BSB

Baanu Sai Sankar Bojja

Data Engineer
5+ Years of Experience
$1.2M+ GMV Recovered