About me
Abhishek Singh
Senior Data Engineer
Senior Data Engineer building cloud data architecture on Azure and GCP. I specialize in event-driven ingestion, Delta lakehouse design, and the serving layers that put fast, reliable analytics in front of the people who need them. Strong on SQL, PySpark, and high-volume data at scale.
Over the last few years at Nuvoretail I've led the move from legacy SQL Server stacks to modern, cloud-native data platforms — designing Data Lake architectures on Azure, orchestrating PySpark pipelines, and standing up cross-cloud analytics on BigQuery. I care about pipelines that are resilient, observable, and boring to operate.
I'm happiest when a messy, petabyte-scale ingestion problem turns into a clean dataset that powers a dashboard people actually trust.
Experience
Senior Data Engineer
Nuvoretail Enlytical Technologies
Designed and built the end-to-end data platform powering enlytical.ai, an AI-enabled ads & analytics platform.
- Built a distributed data-collection platform orchestrating ingestion across 50+ heterogeneous APIs and web sources, landing raw data into a Delta lake for downstream processing.
- Built a centralized job-orchestration service that schedules, dispatches, and monitors crawling and data-processing workloads across 100+ Azure VMs.
- Built a distributed, event-driven ingestion platform on Azure Storage Queue events with asynchronous, fault-tolerant execution — cataloguing and processing terabyte-scale data through medallion layers.
- Implemented a Delta Lake medallion architecture on ADLS with file-level partitioning and scheduled optimization to cut query scans and storage costs.
- Built a config-driven ingestion framework that provisions schemas, transformations, validation rules, and storage layouts for new datasets without code changes — plus a metadata-driven data-reliability suite guarding every layer.
- Built a cross-cloud serving layer via Synapse/Fabric and BigQuery, exposing curated gold datasets to product dashboards and FastAPI services.
- Engineered a low-latency analytics layer with pre-aggregations, query-optimized storage, and in-memory caching — bringing interactive dashboard responses to sub-second.
Data Engineer
Nuvoretail Enlytical Technologies
Owned data architecture, SQL administration, and analytics development.
- Initiated the migration off legacy SQL Server, standing up the first lakehouse foundation — Parquet on ADLS with a Synapse Serverless SQL serving layer — that later grew into the production platform.
- Built and maintained complex SQL- and Spark-based ETL pipelines, including stored procedures and functions, to serve business-critical analytics dashboards.
- Orchestrated data workflows and refresh schedules while managing SQL Server performance, uptime, and processing capacity.
Education
B.Tech, Computer Science & Engineering
Kurukshetra University
2017 — 2021
Certifications
- Microsoft Certified: Azure Data FundamentalsMicrosoft · Mar 2024
- Microsoft Azure Data Engineer Associate (DP-203) Cert PrepLinkedIn · Dec 2025
- Databricks Certified Data Engineer Associate Cert PrepLinkedIn · Dec 2025
- Databricks for Data EngineeringLinkedIn · Dec 2025
- Big Data Analytics with Hadoop and Apache SparkLinkedIn · Dec 2025
- Learning Apache AirflowLinkedIn · Dec 2025
- Advanced SQL for Query Tuning and Performance OptimizationLinkedIn · Dec 2025
- Level Up: Advanced SQLLinkedIn · Dec 2025
- SQL (Advanced)HackerRank · Aug 2022
- SQL (Intermediate)HackerRank · Jun 2022