Microsoft

Experiment Management, Tuning & Debugging

Microsoft

Experiment Management, Tuning & Debugging

 Microsoft

Instructor: Microsoft

What you'll learn

  • Apply LoRA and QLoRA fine-tuning to large language models using Hugging Face PEFT, comparing VRAM usage, throughput, and task performance.

  • Design and execute hyperparameter optimization sweeps on Azure ML using Bayesian sampling, early termination, and MLflow experiment tracking.

  • Diagnose training failure modes, including gradient explosion, overfitting, and normalization errors, using PyTorch Profiler and ablation studies.

  • Build high-throughput data pipelines using WebDataset, LMDB, and Azure ML Data Assets to eliminate I/O bottlenecks and maximize GPU utilization.

Details to know

Shareable certificate

Add to your LinkedIn profile

Assessments

17 assignments¹

AI Graded see disclaimer
Taught in English

See how employees at top companies are mastering in-demand skills

 logos of Petrobras, TATA, Danone, Capgemini, P&G and L'Oreal

Build your Machine Learning expertise

This course is part of the Microsoft Deep Learning Engineering with Azure Professional Certificate
When you enroll in this course, you'll also be enrolled in this Professional Certificate.
  • Learn new concepts from industry experts
  • Gain a foundational understanding of a subject or tool
  • Develop job-relevant skills with hands-on projects
  • Earn a shareable career certificate from Microsoft

There are 9 modules in this course

Establish the foundation of parameter-efficient adaptation. This module explores why full fine-tuning scales poorly on enterprise hardware budgets and mathematically deconstructs how Low-Rank Adaptation (LoRA) compresses weight update tensors into low-rank matrices.

What's included

1 video2 readings1 assignment

Move to practical programming implementation. This module configures, instantiates, and executes quantized parameter-efficient configurations using Hugging Face PEFT, injecting low-rank adapters into attention targets while maintaining strict numerical stability parameters.

What's included

3 readings3 assignments

Automate multi-trial hyperparameter exploration. You will structure cloud infrastructure blueprints to run parallel optimization sweeps using Azure ML SDK v2, defining parameter search spaces and applying early termination rules to drop stagnant training runs.

What's included

1 video3 readings1 assignment

Establish deep observational tracking over your training workloads. You will integrate MLflow tracking hooks into PyTorch scripts, capture essential training artifacts, and evaluate large-scale multi-trial tables inside Azure ML Studio to identify the optimal model configuration.

What's included

1 video2 readings3 assignments

Triage execution anomalies at their source. You will learn to use the PyTorch Profiler and TensorBoard to isolate gradient failures, calculate tensor norm thresholds, and implement gradient clipping to stabilize training steps.

What's included

1 video2 readings1 assignment

Stabilize model generalization limits. You will programmatically configure modern normalization structures (LayerNorm and RMSNorm), apply advanced data-augmentation frameworks via torchvision.transforms.v2, and implement structured ablation workflows to measure your code's robustness against severe overfitting.

What's included

3 readings3 assignments

Bridge the gap between storage layers and computing engines. Learners profile file ingestion patterns to locate hardware starvation issues, configure Azure ML data asset connections, and package individual assets into compressed sharded formats.

What's included

1 video3 readings1 assignment

Unleash the full performance of cloud streaming. Learners configure PyTorch DataLoader parameters for multi-threaded execution, balance input/output modes against local storage boundaries, and verify hardware utilization metrics.

What's included

1 video2 readings3 assignments

Synthesize your model-tuning and experiment-management skills to engineer an automated, memory-efficient fine-tuning pipeline. You will write a Python script that loads a large foundation model in 4-bit precision, configures Low-Rank Adaptation (LoRA) target modules, and instruments a custom training loop with MLflow metric tracking. You will then write the Azure ML SDK v2 configuration code to orchestrate a distributed hyperparameter sweep over your pipeline, utilizing a Bandit early stopping policy to optimize compute cluster resource allocations.

What's included

2 readings1 assignment

Earn a career certificate

Add this credential to your LinkedIn profile, resume, or CV. Share it on social media and in your performance review.

Instructor

 Microsoft
443 Courses2,903,096 learners

Offered by

Microsoft

Why people choose Coursera for their career

Felipe M.

Learner since 2018
"To be able to take courses at my own pace and rhythm has been an amazing experience. I can learn whenever it fits my schedule and mood."

Jennifer J.

Learner since 2020
"I directly applied the concepts and skills I learned from my courses to an exciting new project at work."

Larry W.

Learner since 2021
"When I need courses on topics that my university doesn't offer, Coursera is one of the best places to go."

Chaitanya A.

"Learning isn't just about being better at your job: it's so much more than that. Coursera allows me to learn without limits."

¹ Some assignments in this course are AI-graded. For these assignments, your data will be used in accordance with Coursera's Privacy Notice.