Multi-Source Data Pipeline (Sales)
Full-stack sales data pipeline integrating five source databases through Kafka, Flink, Spark, Airflow and Hadoop to Power BI.
Architecture
Five source databases unified through streaming + batch into a star schema and BI dashboard.
Orchestration
Infrastructure
Overview
A comprehensive sales data pipeline that consolidates five different source databases — PostgreSQL, MySQL, Oracle, SQL Server, and MongoDB — into a unified analytics layer.
It combines streaming and batch processing across Kafka, Flink, and Spark, orchestrated with Airflow, with Hadoop for distributed storage and Docker for packaging. Final insights are delivered in Power BI.
Highlights
- Integrates 5 heterogeneous sources: PostgreSQL, MySQL, Oracle, SQL Server, MongoDB
- Streaming + batch with Kafka, Flink, and Spark
- Orchestrated with Airflow; Hadoop storage; Dockerized
- Business reporting in Power BI
Tech stack
More projects
View all →SMILE Platform (UNDP for Kemenkes)
National immunization logistics system in production across all of Indonesia — streaming CDC + batch analytics platform.
LidValid — Unified Data Validation Platform
Full-stack data validation platform with Tiered Validation — cheap aggregate checks across every table, then precise row-level checks only where they fail.
Dataklin — AI-Powered Data Quality & Entity Resolution
Data quality and entity-resolution platform that profiles, validates, deduplicates, and standardizes datasets before handoff — with LLM-generated rules.