Deploying Deep Learning: Quantization, Serving, and Edge AI
Completed by User 145347
August 7, 2026
21 hours (approximately)
User 145347's account is verified. Coursera certifies their successful completion of Deploying Deep Learning: Quantization, Serving, and Edge AI
What you will learn
Apply INT4/INT8 quantization (AWQ, GPTQ, GGUF) to compress LLMs and vision models for production
Deploy high-throughput inference servers using vLLM's PagedAttention and NVIDIA Triton
Run optimized LLMs on CPU and edge devices using ONNX Runtime and Llama.cpp
Build, benchmark, and containerize a production-ready inference API with Docker
Skills you will gain
- Category: Application Deployment
- Category: LLM Application
- Category: Containerization
- Category: Fine-tuning
- Category: Model Evaluation
- Category: Memory Management
- Category: Docker (Software)
- Category: Model Deployment
- Category: Model Optimization
- Category: Cloud Deployment
- Category: MLOps (Machine Learning Operations)
- Category: Large Language Modeling

