Edureka

Data Engineering with Open Source Tools Specialization

Edureka

Data Engineering with Open Source Tools Specialization

Build Pipelines That Move Data You Can Trust.

Learn the open-source stack behind them, from SQL and Python through Spark and Kafka

Edureka

Instructor: Edureka

What you'll learn

  • Build and validate relational data workflows using SQL, Python, and PostgreSQL.

  • Process and transform data at scale with PySpark and layered dbt models.

  • Orchestrate batch pipelines with Apache Airflow and manage an Iceberg lakehouse.

  • Build and monitor streaming pipelines with Kafka and Spark Structured Streaming.

Details to know

Shareable certificate

Add to your LinkedIn profile

Taught in English
Recently updated!

September 2026

See how employees at top companies are mastering in-demand skills

 logos of Petrobras, TATA, Danone, Capgemini, P&G and L'Oreal

Advance your subject-matter expertise

  • Learn in-demand skills from university and industry experts
  • Master a subject or tool with hands-on projects
  • Develop a deep understanding of key concepts
  • Earn a career certificate from Edureka

Specialization - 3 course series

What you'll learn

  • Explain core data engineering concepts, lifecycle stages, and the role of batch and real-time processing in modern data systems.

  • Work with CSV, JSON, Parquet, PostgreSQL, and SQL to store, model, query, and analyze structured data.

  • Build reliable file-to-database workflows using data quality checks, cleaning, idempotency, and safe reprocessing.

  • Use Python, Git, GitHub, and command-line tools to organize, manage, and troubleshoot practical data engineering projects.

Skills you'll gain

Category: Data Science
Category: Data Structures
Category: Data Pipelines
Category: Extract, Transform, Load
Category: Data Quality
Category: Database Design
Category: Data Architecture
Category: Data Modeling
Category: Git (Version Control System)
Category: Python Programming
Category: Data Engineering
Category: Data Processing
Category: Data Transformation
Category: Data Management
Category: SQL
Category: Relational Databases
Category: Data Analysis
Category: PostgreSQL
Category: Version Control
Category: Data Validation

What you'll learn

  • Explain distributed data processing, Spark architecture, partitions, shuffles, and execution flow for scalable data workloads.

  • Transform structured data using PySpark and Spark SQL, and build reusable analytical models with dbt Core.

  • Orchestrate reliable batch pipelines using Apache Airflow with task dependencies, scheduling, retries, and monitoring.

  • Build open lakehouse workflows using MinIO and Apache Iceberg with Parquet, snapshots, schema evolution, and compaction.

Skills you'll gain

Category: Data Processing
Category: Apache Spark
Category: Distributed Computing
Category: PySpark
Category: Data Engineering
Category: Data Security
Category: Data Quality
Category: Data Transformation
Category: PostgreSQL
Category: Data Pipelines
Category: SQL
Category: Data Science
Category: Data Lakes
Category: Data Collection
Category: Data Management
Category: Data Architecture
Category: Data Validation
Category: Data Storage
Category: Data Modeling
Category: Apache Airflow

What you'll learn

  • Explain event streaming and Kafka architecture, including brokers, topics, partitions, producers, consumers, offsets, and consumer groups.

  • Build real-time data pipelines using Apache Kafka and process streaming events with Spark Structured Streaming and PySpark.

  • Apply event-time processing, late-data handling, stateful operations, checkpointing, and recovery to reliable streaming workloads.

  • Validate, monitor, and troubleshoot Kafka and Spark pipelines to build reliable, production-ready streaming data workflows.

Skills you'll gain

Category: Data Analysis
Category: Data Management
Category: Data Governance
Category: Data Transformation
Category: Docker (Software)
Category: Data Capture
Category: Data Modeling
Category: Data Science
Category: Data Pipelines
Category: Apache
Category: Data Security
Category: Apache Kafka
Category: Apache Spark
Category: Data Quality
Category: Data Architecture
Category: Data Collection
Category: Python Programming
Category: Data Engineering
Category: PySpark
Category: Data Validation

Earn a career certificate

Add this credential to your LinkedIn profile, resume, or CV. Share it on social media and in your performance review.

Instructor

Edureka
Edureka
258 Courses230,338 learners

Offered by

Edureka

Why people choose Coursera for their career

Felipe M.

Learner since 2018
"To be able to take courses at my own pace and rhythm has been an amazing experience. I can learn whenever it fits my schedule and mood."

Jennifer J.

Learner since 2020
"I directly applied the concepts and skills I learned from my courses to an exciting new project at work."

Larry W.

Learner since 2021
"When I need courses on topics that my university doesn't offer, Coursera is one of the best places to go."

Chaitanya A.

"Learning isn't just about being better at your job: it's so much more than that. Coursera allows me to learn without limits."