Production deep learning engineering demands model compression, accelerated inference, and reliable deployment infrastructure on Azure ML.
You'll apply post-training quantization (INT8/FP16), quantization-aware training, structured pruning, and knowledge distillation to compress models and benchmark accuracy-latency trade-offs. You'll configure CUDA and TensorRT with ONNX Runtime and Hugging Face Optimum for accelerated inference, then deploy containerized models to Azure ML online and batch endpoints with autoscaling and blue/green deployment patterns. By the end of this course, you'll be able to compress production models, accelerate inference, and deploy and manage Azure ML endpoints. You'll apply these skills in an end-to-end engineering project using DeepSpeed, FSDP, MLflow, and Azure ML SDK v2. This course is designed for MLOps and deployment engineers focused on inference latency, memory efficiency, and production endpoints. Hands-on PyTorch and Azure ML experience is expected.














