This Specialization covers practical data engineering using widely adopted open-source technologies.You will begin with relational data modeling, SQL, file processing, and data quality using Python and PostgreSQL. You will then process data at scale with PySpark and Spark SQL, build layered and tested transformation models with dbt Core, orchestrate batch workflows with Apache Airflow, and manage open lakehouse data using Apache Iceberg and MinIO.
You will also build streaming pipelines with Apache Kafka and Spark Structured Streaming, applying event-time processing, windowing, checkpoints, monitoring, and recovery.By the end of this Specialization, you will be able to:
Build and validate relational data workflows.
Process large datasets with PySpark and Spark SQL.
Create layered, tested dbt models.
Orchestrate batch pipelines with Apache Airflow.
Manage open lakehouse data with Apache Iceberg and MinIO.
Build and monitor streaming pipelines with Kafka and Spark Structured Streaming.
It is Designed for aspiring data engineers, analytics engineers, software developers, analysts, and database professionals. Basic Python and SQL knowledge is recommended; prior Spark or Kafka experience is not required.
Enroll now to master open-source data engineering across relational workflows, scalable lakehouse pipelines, and real-time streaming with SQL, Python, Spark, dbt, Airflow, Iceberg, and Kafka.
Applied Learning Project
Across the Specialization, learners will complete hands-on projects that will reflect practical data engineering workflows.
They will begin by building a structured data processing workflow using Python, PostgreSQL, SQL, and common file formats while applying data quality and validation practices.
They will then build an orchestrated batch lakehouse pipeline that will process data using PySpark, create reusable transformations with dbt Core, coordinate workflows using Apache Airflow, and manage analytical data using Apache Iceberg and MinIO.
Learners will then bring together Apache Kafka and Spark Structured Streaming to build a reliable real-time pipeline that will incorporate event processing, windowed aggregations, data validation, checkpoint-based recovery, monitoring, and failure handling.
Together, these projects will provide learners with practical experience in designing and validating batch, lakehouse, and streaming data pipelines using open-source technologies.
















