This course offers a hands-on approach to mastering data engineering using Apache Spark, Delta Lake, and Databricks. By combining these technologies, you will learn how to build robust, scalable data pipelines and implement effective data management strategies in real-world applications. With a focus on performance optimization, data orchestration, and modern data engineering practices, this course provides essential skills for professionals working in the data engineering space.

Data Engineering with Databricks Cookbook

What you'll learn
Implement Apache Spark for efficient data ingestion and transformation
Optimize performance of Spark and Delta Lake for scalable data solutions.
Build and orchestrate data pipelines using Databricks workflows and Delta Live Tables.
Skills you'll gain
Tools you'll learn
Details to know

Add to your LinkedIn profile
11 assignments
June 2026
See how employees at top companies are mastering in-demand skills

There are 11 modules in this course
This module introduces practical techniques for ingesting and extracting data from various formats such as CSV, JSON, and XML using Apache Spark. Learners will explore common challenges, data transformation functions, and methods for handling nested and complex data structures. By the end, participants will be equipped to efficiently process and manipulate diverse data sources in Spark.
What's included
1 video8 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
8 readings•Total 40 minutes
- Introduction•4 minutes
- Common Issues Faced While Working with CSV Data•4 minutes
- Reading JSON Data with Apache Spark•5 minutes
- The Flatten() and Collect_list() Functions•6 minutes
- Parsing XML Data with Apache Spark•4 minutes
- Working with Nested Data Structures in Apache Spark•5 minutes
- The Map Keys and Map Values Functions•6 minutes
- Using the regexp_extract() Function•6 minutes
1 assignment•Total 16 minutes
- Data Ingestion and Extraction with Apache Spark•16 minutes
This module introduces learners to essential data manipulation techniques using Apache Spark and PySpark, including filtering, joining, aggregating, and handling null values in large datasets. Learners will explore both standard and advanced operations such as approximate aggregations and nested window functions to efficiently process and analyze data. By the end, participants will be equipped to transform and manage data at scale using Spark's distributed computing capabilities.
What's included
1 video7 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
7 readings•Total 34 minutes
- Introduction•6 minutes
- Filtering Data with Apache Spark•5 minutes
- Performing Joins with Apache Spark•5 minutes
- Performing Aggregations with Apache Spark•4 minutes
- Approximate Aggregations•6 minutes
- Nested Window Functions•5 minutes
- Handling Null Values with Apache Spark•3 minutes
1 assignment•Total 16 minutes
- Mastering Data Processing in Apache Spark•16 minutes
This module introduces the core concepts and practical skills needed to manage data using Delta Lake, an open-source storage layer for lakehouse architectures. Learners will explore reading and merging data, implementing change data capture, optimizing tables, and leveraging versioning and time travel features to ensure data integrity and performance. Hands-on exercises will reinforce best practices for handling big data workloads with Delta Lake in Python.
What's included
1 video6 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
6 readings•Total 37 minutes
- Introduction•7 minutes
- Reading a Delta Lake Table•5 minutes
- Merging Data into Delta Tables•7 minutes
- Change Data Capture in Delta Lake•5 minutes
- Optimizing Delta Lake Tables•6 minutes
- Versioning and Time Travel for Delta Lake Tables•7 minutes
1 assignment•Total 16 minutes
- Mastering Delta Lake Data Management•16 minutes
This module introduces the fundamentals of processing real-time data streams using Apache Spark Structured Streaming. Learners will explore how to ingest data from sources like Apache Kafka, apply transformations and filters, configure checkpoints and triggers, and perform windowed aggregations for robust stream processing applications.
What's included
1 video6 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
6 readings•Total 42 minutes
- Introduction•9 minutes
- Reading Data from Real-Time Sources, Such as Apache Kafka, with Apache Spark Structured Streaming•7 minutes
- Defining Transformations and Filters on a Streaming DataFrame•4 minutes
- Configuring Checkpoints for Structured Streaming in Apache Spark•6 minutes
- Configuring Triggers for Structured Streaming in Apache Spark•6 minutes
- Applying Window Aggregations to Streaming Data with Apache Spark Structured Streaming•10 minutes
1 assignment•Total 16 minutes
- Exploring Streaming Data Processing with Apache Spark•16 minutes
This module explores real-time data processing using Apache Spark Structured Streaming and Delta Lake. Learners will discover techniques for idempotent stream writing, merging change data capture events, joining streaming and static datasets, and monitoring streaming queries. Practical recipes and examples will help you build robust, scalable streaming data pipelines.
What's included
1 video6 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
6 readings•Total 38 minutes
- Introduction•8 minutes
- Idempotent Stream Writing with Delta Lake and Apache Spark Structured Streaming•6 minutes
- Merging or Applying Change Data Capture on Apache Spark Structured Streaming and Delta Lake•6 minutes
- Joining Streaming Data with Static Data in Apache Spark Structured Streaming and Delta Lake•5 minutes
- Joining Streaming Data with Streaming Data in Apache Spark Structured Streaming and Delta Lake•6 minutes
- Monitoring Real-Time Data Processing with Apache Spark Structured Streaming•7 minutes
1 assignment•Total 16 minutes
- Streaming Data Processing Fundamentals•16 minutes
This module explores advanced techniques for optimizing Apache Spark applications, focusing on improving performance and resource efficiency. Learners will discover strategies such as minimizing data shuffling, handling data skew, leveraging broadcast variables, and optimizing partitioning and join operations. Practical guidance on caching and persistence will also be provided to help accelerate data processing workflows.
What's included
1 video7 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
7 readings•Total 46 minutes
- Introduction•5 minutes
- Using Broadcast Variables•5 minutes
- Optimizing Spark Jobs by Minimizing Data Shuffling•6 minutes
- Avoiding Data Skew•8 minutes
- Caching and Persistence•5 minutes
- Partitioning and Repartitioning•8 minutes
- Optimizing Join Strategies•9 minutes
1 assignment•Total 16 minutes
- Mastering Spark Performance Tuning•16 minutes
This module explores advanced techniques to enhance query performance in Delta Lake, including data partitioning, Z-ordering, data skipping, and compression strategies. Learners will gain practical skills to optimize storage and reduce I/O costs for large-scale data processing.
What's included
1 video4 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
4 readings•Total 23 minutes
- Introduction•8 minutes
- Organizing Data with Z-ordering for Efficient Query Execution•6 minutes
- Skipping Data for Faster Query Execution•4 minutes
- Reducing Delta Lake Table Size and I/O Cost with Compression•5 minutes
1 assignment•Total 16 minutes
- Performance Tuning in Delta Lake•16 minutes
This module introduces learners to automating and managing data pipelines using Databricks Workflows. You will explore how to configure, monitor, and parameterize workflows, implement conditional branching, and trigger jobs based on external events such as file arrivals. By the end, you'll be equipped to orchestrate robust data processing tasks on the Databricks platform.
What's included
1 video5 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
5 readings•Total 30 minutes
- Introduction•8 minutes
- Running and Managing Databricks Workflows•3 minutes
- Passing Task and Job Parameters Within a Databricks Workflow•5 minutes
- Conditional Branching in Databricks Workflows•6 minutes
- Triggering Jobs Based on File Arrival•8 minutes
1 assignment•Total 16 minutes
- Mastering Databricks Workflow Orchestration•16 minutes
This module guides learners through building robust data pipelines using Delta Live Tables on Databricks. You will explore techniques for ingesting and transforming streaming data, enforcing data quality, quarantining invalid records, monitoring pipeline health, deploying with asset bundles, and implementing change data capture (CDC). By the end, you'll be equipped to create scalable, reliable pipelines for real-time analytics.
What's included
1 video7 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
7 readings•Total 39 minutes
- Introduction•6 minutes
- Building a Data Pipeline with Delta Live Tables on Databricks•4 minutes
- Implementing Data Quality and Validation Rules with Delta Live Tables in Databricks•6 minutes
- Quarantining Bad Data with Delta Live Tables in Databricks•4 minutes
- Monitoring Delta Live Tables Pipelines•4 minutes
- Deploying Delta Live Tables Pipelines with Databricks Asset Bundles•9 minutes
- Applying Changes (CDC) to Delta Tables with Delta Live Tables•6 minutes
1 assignment•Total 16 minutes
- Data Pipeline Fundamentals with Delta Live Tables•16 minutes
This module introduces the core features of Databricks Unity Catalog for managing data governance in a lakehouse environment. Learners will explore catalog creation, fine-grained access controls, metadata management, data lineage, and system table querying to ensure secure and compliant data operations. Practical exercises demonstrate how to implement row filters, column masks, and leverage the Unity Catalog UI for effective data stewardship.
What's included
1 video9 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
9 readings•Total 40 minutes
- Introduction•7 minutes
- Creating a Catalog•4 minutes
- Defining and Applying Fine-Grained Access Control Policies Using Unity Catalog•5 minutes
- Tagging, Commenting, and Capturing Metadata About Data and AI Assets Using Databricks Unity Catalog•5 minutes
- Using the Unity Catalog UI•4 minutes
- Apply Row Filters•4 minutes
- Apply Column Masks•3 minutes
- Using Unity Catalogs Lineage Data for Debugging, Root Cause Analysis, and Impact Assessment•4 minutes
- Accessing and Querying System Tables Using Unity Catalog•4 minutes
1 assignment•Total 16 minutes
- Data Governance with Unity Catalog•16 minutes
This module explores practical strategies for implementing DataOps and DevOps workflows on the Databricks platform. Learners will discover how to automate tasks using the Databricks CLI, streamline development with the VSCode extension, manage infrastructure with Databricks Asset Bundles, and integrate CI/CD pipelines using GitHub Actions. By the end, participants will be equipped to enhance data and software development efficiency through automation and best practices.
What's included
1 video5 readings1 assignment
1 video•Total 1 minute
- Overview•1 minute
5 readings•Total 35 minutes
- Introduction•9 minutes
- Automating Tasks by Using the Databricks CLI•6 minutes
- Using the Databricks VSCode Extension for Local Development and Testing•4 minutes
- Using Databricks Asset Bundles (DABs)•8 minutes
- Leveraging GitHub Actions with Databricks Asset Bundles (DABs)•8 minutes
1 assignment•Total 16 minutes
- DataOps and DevOps Implementation on Databricks•16 minutes
Instructor

Offered by
Why people choose Coursera for their career

Felipe M.

Jennifer J.

Larry W.

