The Ultimate Hands-On Hadoop

The Ultimate Hands-On Hadoop

This course is part of Big Data Foundations with Hadoop and Spark Specialization

Instructor: Packt - Course Instructors

Access provided by Marlabs

12 modules

Gain insight into a topic and learn the fundamentals.

Intermediate level

Recommended experience

2 weeks to complete

at 10 hours a week

Flexible schedule

Learn at your own pace

12 modules

Gain insight into a topic and learn the fundamentals.

Intermediate level

Recommended experience

2 weeks to complete

at 10 hours a week

Flexible schedule

Learn at your own pace

What you'll learn

Remember Hadoop setup and configuration steps.
Understand the Hadoop ecosystem, including HDFS, MapReduce, and YARN.
Apply queries using Pig, Hive, and Spark.
Evaluate Hadoop cluster performance and optimize it.

Skills you'll gain

Details to know

Shareable certificate

Add to your LinkedIn profile

Assessments

5 assignments

Taught in English

See how employees at top companies are mastering in-demand skills

Learn more about Coursera for Business

logos of Petrobras, TATA, Danone, Capgemini, P&G and L'Oreal

Build your subject-matter expertise

This course is part of the Big Data Foundations with Hadoop and Spark Specialization

When you enroll in this course, you'll also be enrolled in this Specialization.

Learn new concepts from industry experts
Gain a foundational understanding of a subject or tool
Develop job-relevant skills with hands-on projects
Earn a shareable career certificate

There are 12 modules in this course

Updated in May 2025.

This course now features Coursera Coach — your interactive learning companion that helps you test your knowledge, challenge assumptions, and deepen your understanding as you progress. Build a strong, hands-on foundation in Hadoop and big data processing with this comprehensive course designed for data engineers, developers, and IT professionals. From installation to advanced analytics, you’ll learn how to work confidently with Hadoop’s ecosystem and design scalable solutions for real-world data challenges. You’ll begin by installing the Hortonworks Data Platform (HDP) Sandbox on your local machine, giving you an isolated environment to explore Hadoop’s core components. Through guided exercises, you’ll work with the Hadoop Distributed File System (HDFS) and build your understanding of MapReduce, learning how large-scale distributed processing works behind the scenes. As you progress, you’ll move into advanced Hadoop programming with Pig, Hive, and Spark. You’ll write complex queries, analyze large datasets, and work with real-world data to build scalable data workflows. You’ll also explore machine learning with Spark MLLib, giving you a practical introduction to distributed ML techniques. In the final modules, you’ll learn how to manage and optimize Hadoop clusters using YARN, ZooKeeper, Oozie, and Kafka. You’ll practice feeding data into your cluster, orchestrating workflows, managing resources, and analyzing streaming data in real time — essential skills for production-grade environments. By the end of this course, you will have: - Installed and configured the Hortonworks Sandbox for Hadoop development. - Worked with HDFS, MapReduce, and Hadoop’s core data processing concepts. - Written queries and pipelines using Pig, Hive, and Spark. - Performed distributed machine learning with Spark MLLib. - Integrated relational and non-relational data sources with Hadoop. - Managed clusters and streaming workflows with YARN, ZooKeeper, Oozie, and Kafka. - Gained the confidence to design and implement Hadoop-based data solutions. This course is ideal for data engineers, developers, and IT professionals with basic programming or data management experience. Familiarity with Java, SQL, or the Linux command line is helpful but not required.

In this module, we will dive into the world of Hadoop, starting with its installation and setup using the Hortonworks Data Platform Sandbox. You'll explore the key buzzwords and technologies that make up the Hadoop ecosystem, learn about the historical context and impact of the Hortonworks and Cloudera merger, and begin working with real data to get a feel for Hadoop's capabilities.

What's included

4 videos2 readings

In this module, we will explore the core components of Hadoop: the Hadoop Distributed File System (HDFS) and MapReduce. You'll learn how HDFS reliably stores massive data sets across a cluster and how MapReduce enables distributed data processing. Through hands-on activities, you'll import datasets, set up a MapReduce environment, and write scripts to analyze data, including breaking down movie ratings and ranking movies by popularity.

What's included

10 videos1 plugin

10 videosTotal 94 minutes

Hadoop Distributed File System (HDFS): What it is and How it Works13 minutes
Installing the MovieLens Dataset6 minutes
Activity - Installing the MovieLens Dataset into Hadoop's Distributed File System (HDFS) using the Command Line7 minutes
MapReduce: What it is and How it Works10 minutes
How MapReduce Distributes Processing12 minutes
MapReduce Example: Breaking Down the Movie Ratings by Rating Score11 minutes
Activity - Installing Python, MRJob, and Nano7 minutes
Activity - Coding Up and Running the Ratings Histogram MapReduce Job7 minutes
Exercise - Ranking Movies by Their Popularity7 minutes
Activity - Checking Results8 minutes

1 pluginTotal 20 minutes

Exploring HDFS Core Concepts and Architecture20 minutes

In this module, we will delve into Pig, a high-level scripting language that simplifies Hadoop programming. You'll start by exploring the Ambari web-based UI, which makes working with Pig more accessible. The module includes practical examples and activities, such as finding the oldest five-star movies and identifying the most-rated one-star movies using Pig scripts. You'll also learn about the capabilities of Pig Latin and test your skills through challenges and result comparisons.

What's included

7 videos1 assignment1 plugin

7 videosTotal 56 minutes

Introducing Ambari9 minutes
Introducing the Pig6 minutes
Example - Finding the Oldest Movie with Five-Star Rating Using the Pig15 minutes
Activity - Finding the Old Five-Star Movies with Pig9 minutes
More Pig Latin7 minutes
Exercise - Finding the Most-Rated One-Star Movie1 minute
Pig Challenge - Comparing Results5 minutes

1 assignmentTotal 15 minutes

Assessment 115 minutes

1 pluginTotal 20 minutes

Getting to Know Apache Ambari20 minutes

In this module, we will explore the power of Apache Spark, a key technology in the Hadoop ecosystem known for its speed and versatility. You’ll start by understanding why Spark is a game-changer in big data. The module will cover Resilient Distributed Datasets (RDDs) and Datasets, showing you how to use them to analyze movie ratings data. You'll also delve into Spark's machine learning library (MLLib) to create a movie recommendation system. Through hands-on activities, you'll practice writing Spark scripts and refining your data analysis skills.

What's included

8 videos1 plugin

8 videosTotal 74 minutes

Why Spark?10 minutes
The Resilient Distributed Datasets (RDD)10 minutes
Activity - Finding the Movie with the Lowest Average Rating with the Resilient Distributed Datasets (RDD)15 minutes
Datasets and Spark 2.06 minutes
Activity - Finding the movie with the Lowest Average Rating with DataFrames10 minutes
Activity - Recommending a Movie with Spark's Machine Learning Library (MLLib)12 minutes
Exercise - Filtering the Lowest-Rated Movies by Number of Ratings2 minutes
Activity - Checking Results6 minutes

1 pluginTotal 20 minutes

Understanding Apache Spark RDDs and Transformations20 minutes

In this module, we will explore the integration of relational datastores with Hadoop, focusing on Apache Hive and MySQL. You'll start by learning how Hive enables SQL queries on data within HDFS, followed by hands-on activities to find popular and highly-rated movies using Hive. The module also covers the installation and integration of MySQL with Hadoop, using Sqoop to seamlessly transfer data between MySQL and Hadoop's HDFS/Hive. Through practical exercises, you'll gain proficiency in managing and querying relational data within the Hadoop ecosystem.

What's included

9 videos1 plugin

9 videosTotal 63 minutes

What is Hive?6 minutes
Activity - Using Hive to Find the Most Popular Movie10 minutes
How Hive Works?9 minutes
Exercise - Using Hive to Find the Movie with the Highest Average Rating1 minute
Comparing Solutions4 minutes
Integrating MySQL with Hadoop8 minutes
Activity - Installing MySQL and Importing Movie Data7 minutes
Activity - Using Sqoop to Import Data from MySQL to HFDS/Hive7 minutes
Activity - Using Sqoop to Export Data from Hadoop to MySQL7 minutes

1 pluginTotal 15 minutes

Integrating Hive with Hadoop for SQL Queries15 minutes

In this module, we will explore the use of non-relational (NoSQL) data stores within the Hadoop ecosystem. You'll learn why NoSQL databases are crucial for scalability and efficiency, and dive into specific technologies like HBase, Cassandra, and MongoDB. Through a series of activities, you'll practice importing data into HBase, integrating it with Pig, and using Cassandra and MongoDB alongside Spark. The module concludes with exercises to help you choose the most suitable NoSQL database for different scenarios, empowering you to make informed decisions in big data management.

What's included

12 videos1 assignment1 plugin

12 videosTotal 147 minutes

Why NoSQL?13 minutes
What is HBase?12 minutes
Activity - Importing Movie Ratings into HBase13 minutes
Activity - Using HBase with Pig to Import Data at Scale11 minutes
Cassandra - Overview14 minutes
Activity - Installing Cassandra11 minutes
Activity - Writing Spark Output into Cassandra11 minutes
MongoDB - Overview17 minutes
Activity - Installing MongoDB and Integrating Spark with MongoDB12 minutes
Activity - Using the MongoDB Shell7 minutes
Choosing Database Technology15 minutes
Exercise - Choosing a Database for a Given Problem5 minutes

1 assignmentTotal 15 minutes

Assessment 215 minutes

1 pluginTotal 25 minutes

Understanding NoSQL with MongoDB25 minutes

In this module, we will focus on interactive querying tools that allow you to quickly access and analyze big data across multiple sources. You'll explore technologies like Drill, Phoenix, and Presto, learning how each one solves specific challenges in querying large datasets. The module includes hands-on activities where you'll set up these tools, execute queries that span across databases such as MongoDB, Hive, HBase, and Cassandra, and integrate these tools with other Hadoop ecosystem components. By the end of this module, you'll be equipped to perform efficient, real-time data analysis across varied data stores.

What's included

9 videos1 plugin

9 videosTotal 81 minutes

Overview of Drill7 minutes
Activity - Setting Up Drill10 minutes
Activity - Querying Across Multiple Databases with Drill7 minutes
Overview of Phoenix8 minutes
Activity - Installing Phoenix and Querying HBase7 minutes
Activity - Integrating Phoenix with the Pig11 minutes
Overview of Presto6 minutes
Activity - Installing Presto and Querying Hive12 minutes
Activity - Querying Both Cassandra and Hive Using Presto9 minutes

1 pluginTotal 20 minutes

Exploring Hadoop Query Engines: Apache Drill, Phoenix, and Presto20 minutes

In this module, we will explore the critical components involved in managing a Hadoop cluster. You'll learn about YARN's resource management capabilities, how Tez optimizes task execution using Directed Acyclic Graphs, and the differences between Mesos and YARN. We'll dive into ZooKeeper for maintaining reliable operations and Oozie for orchestrating complex workflows. Hands-on activities will guide you through setting up and using Zeppelin for interactive data analysis and using Hue for a more user-friendly interface. The module also touches on other noteworthy technologies like Chukwa and Ganglia, providing a comprehensive understanding of cluster management in Hadoop.

What's included

13 videos1 plugin

13 videosTotal 119 minutes

Yet Another Resource Negotiator (YARN)10 minutes
Tez4 minutes
Activity - Using Hive on Tez and Measuring the Performance Benefit8 minutes
Mesos7 minutes
ZooKeeper13 minutes
Activity - Simulating a Failing Master with ZooKeeper6 minutes
Oozie11 minutes
Activity - Setting Up a Simple Oozie Workflow16 minutes
Zeppelin - Overview5 minutes
Hands-On with Zeppelin for Spark and MovieLens Analysis12 minutes
SQL and Data Visualization in Zeppelin: MovieLens Analysis with Spark9 minutes
Hue - Overview8 minutes
Other Technologies Worth Mentioning4 minutes

1 pluginTotal 20 minutes

Understanding Apache YARN20 minutes

In this module, we will explore the essential tools for feeding data into your Hadoop cluster, focusing on Kafka and Flume. You'll learn how Kafka supports scalable and reliable data collection across a cluster and how to set it up to publish and consume data. Additionally, you'll discover how Flume's architecture differs from Kafka and how to use it for real-time data ingestion. Through hands-on activities, you'll configure Kafka to monitor Apache logs and Flume to watch directories, publishing incoming data into HDFS. These skills will help you manage and process streaming data effectively in your Hadoop environment.

What's included

6 videos1 assignment1 plugin

6 videosTotal 54 minutes

Kafka9 minutes
Activity - Setting Up Kafka and Publishing Data7 minutes
Activity - Publishing Web Logs with Kafka10 minutes
Flume10 minutes
Activity - Setting up Flume and Publishing Logs7 minutes
Activity - Setting Up Flume to Monitor a Directory and Store its Data in Hadoop Distributed File System (HDFS)9 minutes

1 assignmentTotal 15 minutes

Assessment 315 minutes

1 pluginTotal 20 minutes

Publishing Data to a Big Data Cluster with Kafka20 minutes

In this module, we will focus on analyzing streams of data using real-time processing frameworks such as Spark Streaming, Apache Storm, and Flink. You’ll start by learning how Spark Streaming processes micro-batches of data in real-time and participate in activities that include analyzing web logs streamed by Flume. The module then introduces Apache Storm and Flink, providing hands-on exercises to implement word count applications with these tools. By the end of this module, you will be able to build continuous applications that efficiently process and analyze streaming data.

What's included

8 videos1 plugin

8 videosTotal 76 minutes

Spark Streaming: Introduction14 minutes
Activity - Analyzing Web Logs Published with Flume using Spark Streaming14 minutes
Exercise - Monitor Flume-Published Logs for Errors in Real Time2 minutes
Exercise Solution: Aggregating the Hypertext Transfer Protocol (HTTP) Access Codes with Spark Streaming4 minutes
Apache Storm: Introduction9 minutes
Activity - Counting Words with Storm14 minutes
Flink: Overview6 minutes
Activity - Counting Words with Flink10 minutes

1 pluginTotal 15 minutes

Real-Time Data Processing with Spark Streaming15 minutes

In this module, we will focus on designing and implementing real-world systems using a combination of Hadoop ecosystem tools. You'll start by exploring additional technologies like Impala, NiFi, and AWS Kinesis, learning how they fit into broader Hadoop-based solutions. The module then guides you through the process of understanding system requirements and designing applications that consume and analyze large-scale data, such as web server logs or movie recommendations. By the end of this module, you’ll be equipped to design and build complex, efficient, and scalable data systems tailored to specific business needs.

What's included

7 videos1 assignment1 plugin

7 videosTotal 52 minutes

The Best of the Rest9 minutes
Review: How the Pieces Fit Together?6 minutes
Understanding Your Requirements8 minutes
Sample Application: Consuming Web Server Logs and Keeping Track of Top-Sellers10 minutes
Sample Application: Serving Movie Recommendations to a Website11 minutes
Exercise - Designing a System to Report Web Sessions Per Day2 minutes
Exercise Solution: Designing a System to Count Daily Sessions4 minutes

1 assignmentTotal 15 minutes

Assessment 415 minutes

1 pluginTotal 20 minutes

Navigating the Distributed Computing Landscape20 minutes

In this final module, we will provide you with a selection of books, online resources, and tools recommended by the author to further your knowledge of Hadoop and related technologies. This module serves as a guide for continued learning, offering you the means to stay updated with the latest developments in the Hadoop ecosystem and expand your skills beyond this course.