Three years designing batch and real-time ETL/ELT across telecom and finance — from subscriber usage records to fraud detection and Customer 360 risk datasets, built to survive audits, SLAs, and scale.
SV / Data Engineer
I build pipelines that data teams can trust — the kind that hold up under regulatory scrutiny and peak-hour load alike. My work sits at the intersection of telecom-scale batch processing and finance-grade compliance analytics, turning raw usage logs and transaction records into datasets that risk, fraud, and reporting teams depend on every day.
Built a real-time ingestion pipeline using Kafka to stream clickstream events into Spark Structured Streaming, applying watermarking and windowed aggregations, with results persisted to S3 and Redshift for analytics.
Built a batch ETL pipeline extracting data from a public API into an AWS S3 data lake using PySpark on EMR, modeled a star-schema warehouse in Redshift, orchestrated with Airflow, and visualized in Tableau.