All work
07Data Engineering

Spark Pipeline (Sales)

Streaming & batch ETL with Kafka + Spark, stored in Delta Lake, scheduled with Dagster.

Architecture

Streaming + batch ETL in Spark, stored in Delta Lake and served from PostgreSQL, scheduled with Dagster.

Sales DataCSV producer
Apache KafkaEvent streaming
Apache SparkStreaming + batch ETL
Delta LakeStorage layer
PostgreSQLServing DB

Orchestration

Dagster

Overview

A sales pipeline using Apache Spark for both streaming and batch ETL.

Data is ingested through Kafka, processed in Spark, stored in Delta Lake with PostgreSQL alongside, and the whole thing is scheduled and monitored with Dagster.

Highlights

  • Streaming + batch ETL in Apache Spark
  • Delta Lake storage layer
  • Scheduled and observed with Dagster

Tech stack

KafkaApache SparkDelta LakePostgreSQLDagster