Abhishek.

About me

Abhishek Singh

Senior Data Engineer

Senior Data Engineer building cloud data architecture on Azure and GCP. I specialize in event-driven ingestion, Delta lakehouse design, and the serving layers that put fast, reliable analytics in front of the people who need them. Strong on SQL, PySpark, and high-volume data at scale.

Over the last few years at Nuvoretail I've led the move from legacy SQL Server stacks to modern, cloud-native data platforms — designing Data Lake architectures on Azure, orchestrating PySpark pipelines, and standing up cross-cloud analytics on BigQuery. I care about pipelines that are resilient, observable, and boring to operate.

I'm happiest when a messy, petabyte-scale ingestion problem turns into a clean dataset that powers a dashboard people actually trust.

Experience

Senior Data Engineer

Nuvoretail Enlytical Technologies

Jun 2025 — Present

Designed and built the end-to-end data platform powering enlytical.ai, an AI-enabled ads & analytics platform.

  • Built a distributed data-collection platform orchestrating ingestion across 50+ heterogeneous APIs and web sources, landing raw data into a Delta lake for downstream processing.
  • Built a centralized job-orchestration service that schedules, dispatches, and monitors crawling and data-processing workloads across 100+ Azure VMs.
  • Built a distributed, event-driven ingestion platform on Azure Storage Queue events with asynchronous, fault-tolerant execution — cataloguing and processing terabyte-scale data through medallion layers.
  • Implemented a Delta Lake medallion architecture on ADLS with file-level partitioning and scheduled optimization to cut query scans and storage costs.
  • Built a config-driven ingestion framework that provisions schemas, transformations, validation rules, and storage layouts for new datasets without code changes — plus a metadata-driven data-reliability suite guarding every layer.
  • Built a cross-cloud serving layer via Synapse/Fabric and BigQuery, exposing curated gold datasets to product dashboards and FastAPI services.
  • Engineered a low-latency analytics layer with pre-aggregations, query-optimized storage, and in-memory caching — bringing interactive dashboard responses to sub-second.

Data Engineer

Nuvoretail Enlytical Technologies

Aug 2022 — Jun 2025

Owned data architecture, SQL administration, and analytics development.

  • Initiated the migration off legacy SQL Server, standing up the first lakehouse foundation — Parquet on ADLS with a Synapse Serverless SQL serving layer — that later grew into the production platform.
  • Built and maintained complex SQL- and Spark-based ETL pipelines, including stored procedures and functions, to serve business-critical analytics dashboards.
  • Orchestrated data workflows and refresh schedules while managing SQL Server performance, uptime, and processing capacity.

Education

B.Tech, Computer Science & Engineering

Kurukshetra University

2017 — 2021

Certifications

  • Microsoft Certified: Azure Data FundamentalsMicrosoft · Mar 2024
  • Microsoft Azure Data Engineer Associate (DP-203) Cert PrepLinkedIn · Dec 2025
  • Databricks Certified Data Engineer Associate Cert PrepLinkedIn · Dec 2025
  • Databricks for Data EngineeringLinkedIn · Dec 2025
  • Big Data Analytics with Hadoop and Apache SparkLinkedIn · Dec 2025
  • Learning Apache AirflowLinkedIn · Dec 2025
  • Advanced SQL for Query Tuning and Performance OptimizationLinkedIn · Dec 2025
  • Level Up: Advanced SQLLinkedIn · Dec 2025
  • SQL (Advanced)HackerRank · Aug 2022
  • SQL (Intermediate)HackerRank · Jun 2022