Background
Sections
IntroductionModule 01 β€” Tensors 🧊01 Β· Creating Tensors 🧊02 Β· Indexing & Reshaping πŸ”ͺ03 Β· Tensor Math βž—04 Β· Device Placement πŸ–₯️⚑Module 02 β€” Autograd βš™οΈ01 Β· The Computational Graph πŸ•ΈοΈ02 Β· Backward Pass & Gradients ⬅️03 Β· Turning Autograd Off πŸ›‘04 Β· Gradient Gotchas πŸͺ€Module 03 β€” Neural Networks01 Β· The `nn.Module` Basics 🧱02 Β· Common Layers 🧩03 Β· Building a Network πŸ—οΈ04 Β· Inspecting Models πŸ”Module 04 β€” Data Handling πŸ—‚οΈ01 Β· Dataset Basics πŸ“‡02 Β· The DataLoader 🚚03 Β· Transforms 🎨04 Β· Splits & Built-in Datasets βœ‚οΈModule 05 β€” The Training Loop πŸ”01 Β· Loss Functions 🎯02 Β· Optimizers βš™οΈ03 Β· The Training Loop πŸ”04 Β· Evaluation & Metrics πŸ“ŠModule 06 β€” Saving & Loading πŸ’Ύ01 Β· `state_dict` Basics πŸ’Ύ02 Β· Checkpoints & Resuming ⏸️03 Β· Loading for Inference πŸš€04 Β· Best Model & Early Stopping πŸ…Module 07 β€” Computer Vision πŸ‘οΈ01 Β· Convolutions πŸ”²02 Β· CNN Architecture πŸ›οΈ03 Β· Transfer Learning πŸ”04 Β· Image Classification Project πŸ§ͺModule 08 β€” NLP & Transformers πŸ’¬01 Β· Text Data & Tokenization πŸ”€02 Β· Embeddings 🧭03 Β· Recurrent Layers & LSTMs πŸ”„04 Β· Intro to Transformers ⚑Module 09 β€” Ecosystem: PyTorch Lightning ⚑01 Β· Why Lightning? πŸ€”02 Β· The LightningModule 🧩03 Β· The Trainer πŸŽ›οΈ04 Β· DataModules & Callbacks 🧰Module 10 β€” Deployment & Optimization πŸš€01 Β· Exporting Models πŸ“¦02 Β· `torch.compile` ⚑03 Β· Inference Optimization πŸͺΆ04 Β· Serving & Next Steps πŸŽ“

Module 07 β€” Computer Vision πŸ‘οΈ

3 min read

Welcome to Phase 3. Here we specialize the general skills you've built into the architecture that made deep learning famous: the convolutional neural network.

Everything through Module 06 was general-purpose β€” tensors, autograd, models, training, saving. Now we apply it to a specific, high-impact domain: images. You could flatten a picture into a long vector and feed it to the nn.Linear layers from Module 03, but you'd throw away the single most important fact about an image: that nearby pixels are related, and that a cat is a cat whether it's in the top-left or bottom-right of the frame. Fully-connected layers are blind to this spatial structure and would need an astronomical number of parameters to cope. Convolutional Neural Networks (CNNs) were designed precisely to exploit spatial structure, and they're the reason computers can now recognize objects, read handwriting, and diagnose scans.

The core idea is the convolution: instead of connecting every pixel to every neuron, a small learnable filter slides across the image looking for a local pattern β€” an edge, a texture, a curve β€” everywhere at once. Stack these layers and the network learns a hierarchy, from simple edges in early layers to whole objects in deep ones. This module teaches the two building blocks (Conv2d and MaxPool2d), how to assemble them into a working CNN, and then the pragmatic superpower of modern vision: transfer learning β€” starting from a model someone else trained on millions of images and adapting it to your problem with a fraction of the data and compute.

This module is split into three sub-modules β€” work through them in order.

πŸ“š Sub-Modules

# Sub-Module What you'll learn
01 Convolutions Why convolutions beat dense layers for images; nn.Conv2d, kernels, channels, feature maps
02 CNN Architecture nn.MaxPool2d and stacking conv β†’ pool β†’ classifier into a complete CNN
03 Transfer Learning Reusing pretrained models by freezing layers and fine-tuning a new head
04 Image Classification Project An end-to-end capstone: data β†’ CNN β†’ train β†’ save best β†’ evaluate, on FashionMNIST

🎯 By the end of this module, you'll be able to...

  • Explain what a convolution does and why CNNs suit images so well.
  • Build a complete CNN from Conv2d, MaxPool2d, and Linear layers.
  • Adapt a pretrained network to your own dataset with transfer learning.
  • Assemble everything into a complete, runnable image-classification project.

βœ… Prerequisites

All of Phases 1 and 2 β€” especially building models (Module 03), the data pipeline with image transforms (Module 04), and the training loop (Module 05).


⬅️ Prev module: 06 Β· Saving & Loading Β· ➑️ Next module: 08 Β· NLP & Transformers