CDR / SMS Loans / Deposits Clickstream SPARK / EMR watermark ⋅ checkpoint join ⋅ aggregate S3 LAKE glue catalog REDSHIFT star schema CUSTOMER 360 risk ⋅ compliance Tableau Power BI
Data Engineer · AWS & Apache Spark

Saiyashwanth
Vottikoti

pipelines that carry the weight of the business

Three years designing batch and real-time ETL/ELT across telecom and finance — from subscriber usage records to fraud detection and Customer 360 risk datasets, built to survive audits, SLAs, and scale.

Saiyashwanth Vottikoti SV / Data Engineer
scroll
01

Summary

I build pipelines that data teams can trust — the kind that hold up under regulatory scrutiny and peak-hour load alike. My work sits at the intersection of telecom-scale batch processing and finance-grade compliance analytics, turning raw usage logs and transaction records into datasets that risk, fraud, and reporting teams depend on every day.

3+
Years in Production Data Engineering
2
Industries — Telecom & Finance
3
Clouds — AWS Primary, Azure & GCP Exposure
02

Experience

Data Engineer FEB 2025 — PRESENT
Wells Fargo, San Francisco, CA — Real-Time Fraud Detection & Customer 360 Analytics
  • Built batch ETL pipelines integrating loans, deposits, and card data into a unified AWS lake.
  • Designed Customer 360 datasets supporting credit risk, exposure, and compliance reporting.
  • Implemented reconciliation and validation checks ensuring audit and regulatory data accuracy.
  • Optimized large-scale Spark joins and aggregations for high-volume financial datasets.
  • Automated data workflows with Apache Airflow to meet regulatory reporting deadlines.
  • Partnered with risk, fraud, and compliance teams to deliver analytics-ready datasets.
ETL Developer DEC 2021 — JUN 2023
Reliance Jio Infocomm Ltd., Hyderabad, India — Real-Time Recharge & Batch Usage Processing
  • Built Spark batch ETL jobs processing CDR, SMS, and internet usage logs for billing systems.
  • Designed daily aggregation pipelines computing subscriber-level usage metrics.
  • Applied watermarking and checkpointing to handle late-arriving, out-of-order events.
  • Optimized Spark transformations and joins to meet SLA timelines on telecom datasets.
  • Managed Apache Airflow workflows automating and monitoring batch processing jobs.
  • Monitored production pipelines, resolved failures, and ensured high availability at peak load.
03

Projects

Kafka · Spark Structured Streaming

Real-Time Streaming Pipeline

Built a real-time ingestion pipeline using Kafka to stream clickstream events into Spark Structured Streaming, applying watermarking and windowed aggregations, with results persisted to S3 and Redshift for analytics.

AWS · EMR · Airflow · Tableau

End-to-End Batch ETL Pipeline on AWS

Built a batch ETL pipeline extracting data from a public API into an AWS S3 data lake using PySpark on EMR, modeled a star-schema warehouse in Redshift, orchestrated with Airflow, and visualized in Tableau.

04

Stack, by Pipeline Stage

01
Ingest
Apache Kafka Structured Streaming Hadoop / HDFS
02
Process
PySpark / Spark SQL Python Schema Evolution Watermarking
03
Store
S3 · Redshift Cassandra · HBase BigQuery Hive Metastore
04
Orchestrate
Apache Airflow Azure Data Factory AWS Glue · Lambda
05
Serve
Tableau Power BI Star / Snowflake Schema
05

Education & Certifications

M.S., Management — Data Analytics
Indiana Wesleyan University, Marion, Indiana
AUG 2023 — APR 2025
B.Tech., Mechanical Engineering
Sri Indu College of Engineering & Technology, Hyderabad
AUG 2017 — JUL 2021