Build practical expertise in modern batch data engineering using Apache Spark, PySpark, dbt Core, Apache Airflow, Apache Iceberg, MinIO, PostgreSQL, and SQL. You will learn how data engineers process large datasets, build reusable transformation models, orchestrate dependent workflows, and manage analytical data using open lakehouse technologies.
You will begin by exploring distributed data processing fundamentals and the architecture of Apache Spark. You will examine drivers, executors, jobs, stages, tasks, lazy evaluation, partitions, shuffles, and execution plans. Through guided demonstrations, you will set up PySpark locally, read CSV, JSON, and Parquet datasets, define schemas, transform DataFrames, and analyse data using Spark SQL. You will then move into transformation engineering with dbt Core, where you will explore ETL and ELT approaches, layered modelling, model grain, and reusable SQL transformations. Using dbt Core with PostgreSQL, you will build staging and intermediate models, create fact and dimension tables, apply tests, generate documentation, inspect lineage, and work with materialization strategies, transformation contracts, and change management. Finally, you will learn how to coordinate batch workflows using Apache Airflow and manage lakehouse data using MinIO and Apache Iceberg. You will work with DAGs, tasks, schedules, dependencies, retries, monitoring, backfills, and catchup. You will also explore object storage, Parquet, Iceberg tables, snapshots, schema evolution, compaction, and small-file management. The course concludes with an orchestrated batch lakehouse project that brings together distributed processing, transformation, orchestration, and open lakehouse storage. By the end of this course, you will be able to: - Explain the fundamentals of distributed data processing and Apache Spark architecture. - Process CSV, JSON, and Parquet datasets using PySpark DataFrames. - Transform, clean, aggregate, and analyse data using PySpark and Spark SQL. - Interpret partitions, shuffles, execution plans, caching, and recomputation. - Explain the role of ETL, ELT, and analytics engineering in modern data workflows. - Build layered transformation models using dbt Core. - Create staging, intermediate, fact, dimension, and mart models. - Apply dbt testing, documentation, lineage, materialization, and transformation contracts. - Orchestrate batch workflows using Apache Airflow. - Manage schedules, dependencies, retries, backfills, catchup, and monitoring. - Create and work with Apache Iceberg tables on MinIO. - Understand snapshots, schema evolution, compaction, and small-file management. - Build an end-to-end orchestrated batch lakehouse pipeline using open-source technologies. Designed for data engineers, aspiring data engineers, analytics engineers, software developers, data analysts, and database professionals, this course prepares you to build scalable, maintainable, and reliable batch data pipelines using modern open-source processing, transformation, orchestration, and lakehouse technologies.















