Background
Sections
IntroductionModule 01 β€” Tensors 🧊01 Β· Creating Tensors 🧊02 Β· Indexing & Reshaping πŸ”ͺ03 Β· Tensor Math βž—04 Β· Device Placement πŸ–₯️⚑Module 02 β€” Autograd βš™οΈ01 Β· The Computational Graph πŸ•ΈοΈ02 Β· Backward Pass & Gradients ⬅️03 Β· Turning Autograd Off πŸ›‘04 Β· Gradient Gotchas πŸͺ€Module 03 β€” Neural Networks01 Β· The `nn.Module` Basics 🧱02 Β· Common Layers 🧩03 Β· Building a Network πŸ—οΈ04 Β· Inspecting Models πŸ”Module 04 β€” Data Handling πŸ—‚οΈ01 Β· Dataset Basics πŸ“‡02 Β· The DataLoader 🚚03 Β· Transforms 🎨04 Β· Splits & Built-in Datasets βœ‚οΈModule 05 β€” The Training Loop πŸ”01 Β· Loss Functions 🎯02 Β· Optimizers βš™οΈ03 Β· The Training Loop πŸ”04 Β· Evaluation & Metrics πŸ“ŠModule 06 β€” Saving & Loading πŸ’Ύ01 Β· `state_dict` Basics πŸ’Ύ02 Β· Checkpoints & Resuming ⏸️03 Β· Loading for Inference πŸš€04 Β· Best Model & Early Stopping πŸ…Module 07 β€” Computer Vision πŸ‘οΈ01 Β· Convolutions πŸ”²02 Β· CNN Architecture πŸ›οΈ03 Β· Transfer Learning πŸ”04 Β· Image Classification Project πŸ§ͺModule 08 β€” NLP & Transformers πŸ’¬01 Β· Text Data & Tokenization πŸ”€02 Β· Embeddings 🧭03 Β· Recurrent Layers & LSTMs πŸ”„04 Β· Intro to Transformers ⚑Module 09 β€” Ecosystem: PyTorch Lightning ⚑01 Β· Why Lightning? πŸ€”02 Β· The LightningModule 🧩03 Β· The Trainer πŸŽ›οΈ04 Β· DataModules & Callbacks 🧰Module 10 β€” Deployment & Optimization πŸš€01 Β· Exporting Models πŸ“¦02 Β· `torch.compile` ⚑03 Β· Inference Optimization πŸͺΆ04 Β· Serving & Next Steps πŸŽ“

Module 02 β€” Autograd βš™οΈ

2 min read

The engine that makes neural networks learn. Once you see how it works, "training a model" stops feeling like magic.

In Module 01 you learned that a tensor is a supercharged NumPy array. Here's the superpower that NumPy can't touch: a PyTorch tensor can remember every operation performed on it, then automatically compute the gradients β€” the derivatives β€” of some final output with respect to every input. That process is called automatic differentiation, and PyTorch's implementation is autograd.

Why does this matter? Training a neural network means repeatedly nudging millions of weights in the direction that reduces error. To know which direction, you need the gradient of the loss with respect to each weight. Computing those derivatives by hand would be hopeless. Autograd does it for you: you write the forward computation as ordinary Python, call .backward(), and every gradient appears. Master this module and the training loop in Phase 2 will feel almost inevitable.

This module is split into three sub-modules β€” work through them in order.

πŸ“š Sub-Modules

# Sub-Module What you'll learn
01 The Computational Graph How PyTorch records operations into a dynamic graph, and the role of requires_grad
02 Backward Pass & Gradients The forward vs. backward pass, calling .backward(), and reading gradients from .grad
03 Turning Autograd Off torch.no_grad(), .detach(), and why inference and evaluation need them
04 Gradient Gotchas Why .grad accumulates (and must be zeroed), leaf vs. non-leaf tensors, and retain_graph

🎯 By the end of this module, you'll be able to...

  • Explain what the computational graph is and how requires_grad puts a tensor on it.
  • Run a forward pass, call .backward(), and interpret the gradients stored in .grad.
  • Correctly disable gradient tracking for inference to save memory and avoid subtle bugs.
  • Understand why gradients accumulate and how zeroing them prepares you for the training loop.

βœ… Prerequisites

Comfort with tensors from Module 01 β€” creating them, doing math, and moving them across devices.


⬅️ Prev module: 01 Β· Tensors Β· ➑️ Next module: 03 Β· Neural Networks