In this Specialization, you’ll build practical skills for designing, deploying, and improving batch and streaming data pipelines on Google Cloud. You’ll work with Dataflow, Apache Beam, Cloud Data Fusion, Pub/Sub, BigQuery, Bigtable, Cloud Storage, and Apache Kafka to move, transform, and process data at scale.
You’ll start with guided Dataflow and Data Fusion labs, then build ETL workflows, transform data with Wrangler, and create batch and real-time pipelines. You’ll also explore streaming architectures, windows, watermarks, triggers, sources and sinks, schemas, state and timers, SQL, DataFrames, and Beam notebooks. Along the way, you’ll address data quality, monitoring, orchestration, IAM, quotas, security, performance, and data locality.
By the end of this Specialization, you’ll be able to:
Design and build scalable batch and streaming pipelines for common data engineering use cases.
Use Dataflow, Data Fusion, Pub/Sub, Kafka, BigQuery, and Bigtable across end-to-end workflows.
Apply data quality, monitoring, security, and performance practices to pipeline operations.
Develop Apache Beam pipelines using streaming concepts, schemas, state, timers, SQL, and notebooks.
Applied Learning Project
You’ll gain applied experience through Google Cloud console labs and guided pipeline builds. Set up the Dataflow Python SDK, run example jobs, and launch template-based Pub/Sub-to-BigQuery streaming pipelines. Build Dataflow ETL pipelines that ingest public data into BigQuery, then create and transform JSON and CSV pipelines in Cloud Data Fusion using Pipeline Studio and Wrangler. Configure batch and real-time Data Fusion pipelines, including Pub/Sub sources. Practice stream processing by reading Pub/Sub messages, applying timestamp windows, and writing to Cloud Storage. Create a Kafka cluster, publish topic data, and process it with a Java Kafka Streams WordCount app. In broader batch and streaming scenarios, design scalable pipelines, apply validation and cleansing, orchestrate and monitor workflows, connect Dataflow to Bigtable, and explore Apache Beam windows, triggers, I/O, schemas, state, timers, SQL, DataFrames, custom containers, IAM, quotas, security, and Beam notebooks.



























