Soumyajit Sarkar Statistical Abstract of a Career, edition 2026

Statistical Abstract Table 3b

Backend Analytics Engineering

Data pipelines and revenue analytics systems

$100M+ revenue tracked

= $5 million

80M+ users analyzed annually

= 2 million users

Designing, building and maintaining the server-side infrastructure behind analytics: ingestion, processing pipelines, storage and real-time reporting.

He architected and maintains an analytics management system that has continuously tracked over $100 million in revenue and analyzed data from over 80 million users annually, with real-time processing, automated reporting and predictive analytics.

Table 3b.1 What the work covers

Real-time data processing
High-throughput, low-latency processing with Apache Kafka, Apache Spark and custom stream processing.
Revenue tracking systems
Accurate, scalable revenue tracking and financial analytics for enterprise applications.
Data pipeline architecture
End-to-end pipelines: ETL and ELT, data validation and automated workflows.
Analytics API development
High-performance analytics APIs and microservices as the data access layer.
Data warehouse design
Warehouse architecture on Snowflake, BigQuery and custom solutions.
Performance optimization
Tuning analytics systems for speed, cost and efficiency.

Table 3b.2 Process

  1. 1

    Requirements analysis

    Analytics requirements, data sources and performance targets, plus an assessment of existing systems.

  2. 2

    Architecture design

    Data flow architecture, component specifications and scalability planning.

  3. 3

    Infrastructure development

    Pipelines, processing systems and storage optimization, built hands-on.

  4. 4

    Integration and testing

    Integration with existing systems; validation of data accuracy, performance and reliability.

  5. 5

    Monitoring and maintenance

    Ongoing monitoring and optimization to keep the system performing.

Table 3b.3 Technology

LanguagesPython, Node.js, Go, Java, Scala
ProcessingApache Spark, Apache Kafka, Apache Airflow, custom streaming
Databases and storagePostgreSQL, MongoDB, Redis, Elasticsearch, ClickHouse, distributed storage
CloudAWS, Google Cloud Platform, Azure, hybrid cloud
Analytics and MLTensorFlow, PyTorch, scikit-learn, custom ML frameworks