Spark Pipeline (Sales)
Streaming & batch ETL with Kafka + Spark, stored in Delta Lake, scheduled with Dagster.
Architecture
Streaming + batch ETL in Spark, stored in Delta Lake and served from PostgreSQL, scheduled with Dagster.
Orchestration
Overview
A sales pipeline using Apache Spark for both streaming and batch ETL.
Data is ingested through Kafka, processed in Spark, stored in Delta Lake with PostgreSQL alongside, and the whole thing is scheduled and monitored with Dagster.
Highlights
- Streaming + batch ETL in Apache Spark
- Delta Lake storage layer
- Scheduled and observed with Dagster
Tech stack
More projects
View all →SMILE Platform (UNDP for Kemenkes)
National immunization logistics system in production across all of Indonesia — streaming CDC + batch analytics platform.
LidValid — Unified Data Validation Platform
Full-stack data validation platform with Tiered Validation — cheap aggregate checks across every table, then precise row-level checks only where they fail.
Dataklin — AI-Powered Data Quality & Entity Resolution
Data quality and entity-resolution platform that profiles, validates, deduplicates, and standardizes datasets before handoff — with LLM-generated rules.