Optimizing Spark and Cloud Data Storage for Analytics
Completed by Rayan Abutawil
April 6, 2026
10 hours (approximately)
Rayan Abutawil's account is verified. Coursera certifies their successful completion of Optimizing Spark and Cloud Data Storage for Analytics
What you will learn
Optimize Spark job performance through strategic partitioning and caching, achieving 30%+ runtime improvements using data access analysis.
Implement transactional data lakes with Delta format, enabling versioning, ACID operations, and schema evolution for reliable datasets.
Provision secure cloud data infrastructure using IAM policies, private networks, and encrypted storage following security best practices.
Evaluate and benchmark storage formats (Parquet, ORC, Avro) to select optimal solutions for analytical workloads and cost efficiency.
Skills you will gain
- Category: Cloud Infrastructure
- Category: Cloud Security
- Category: Cloud Deployment
- Category: Data Storage
- Category: Performance Tuning
- Category: Transaction Processing
- Category: Infrastructure as Code (IaC)
- Category: Data Warehousing
- Category: Cloud Storage
- Category: Data Security
- Category: PySpark
- Category: Data Lakes

