Background
Sections
IntroductionModule 01 β€” Tensors 🧊01 Β· Creating Tensors 🧊02 Β· Indexing & Reshaping πŸ”ͺ03 Β· Tensor Math βž—04 Β· Device Placement πŸ–₯️⚑Module 02 β€” Autograd βš™οΈ01 Β· The Computational Graph πŸ•ΈοΈ02 Β· Backward Pass & Gradients ⬅️03 Β· Turning Autograd Off πŸ›‘04 Β· Gradient Gotchas πŸͺ€Module 03 β€” Neural Networks01 Β· The `nn.Module` Basics 🧱02 Β· Common Layers 🧩03 Β· Building a Network πŸ—οΈ04 Β· Inspecting Models πŸ”Module 04 β€” Data Handling πŸ—‚οΈ01 Β· Dataset Basics πŸ“‡02 Β· The DataLoader 🚚03 Β· Transforms 🎨04 Β· Splits & Built-in Datasets βœ‚οΈModule 05 β€” The Training Loop πŸ”01 Β· Loss Functions 🎯02 Β· Optimizers βš™οΈ03 Β· The Training Loop πŸ”04 Β· Evaluation & Metrics πŸ“ŠModule 06 β€” Saving & Loading πŸ’Ύ01 Β· `state_dict` Basics πŸ’Ύ02 Β· Checkpoints & Resuming ⏸️03 Β· Loading for Inference πŸš€04 Β· Best Model & Early Stopping πŸ…Module 07 β€” Computer Vision πŸ‘οΈ01 Β· Convolutions πŸ”²02 Β· CNN Architecture πŸ›οΈ03 Β· Transfer Learning πŸ”04 Β· Image Classification Project πŸ§ͺModule 08 β€” NLP & Transformers πŸ’¬01 Β· Text Data & Tokenization πŸ”€02 Β· Embeddings 🧭03 Β· Recurrent Layers & LSTMs πŸ”„04 Β· Intro to Transformers ⚑Module 09 β€” Ecosystem: PyTorch Lightning ⚑01 Β· Why Lightning? πŸ€”02 Β· The LightningModule 🧩03 Β· The Trainer πŸŽ›οΈ04 Β· DataModules & Callbacks 🧰Module 10 β€” Deployment & Optimization πŸš€01 Β· Exporting Models πŸ“¦02 Β· `torch.compile` ⚑03 Β· Inference Optimization πŸͺΆ04 Β· Serving & Next Steps πŸŽ“

Module 04 β€” Data Handling πŸ—‚οΈ

3 min read

A model is only as good as the data you feed it β€” and how you feed it. This module is the plumbing that makes real training possible.

Welcome to Phase 2. In Phase 1 you built models and learned how gradients flow; the examples used tiny, hand-made tensors. Real datasets are far messier and far larger: thousands of images that won't all fit in memory, files scattered across folders, labels in a separate CSV, values that need normalizing before a network can learn from them. You need a clean, efficient way to load data sample-by-sample, transform it, group it into batches, and shuffle it every epoch. That's exactly what torch.utils.data provides.

PyTorch splits this responsibility into two cooperating pieces, and understanding the division is the key to this module. A Dataset answers a single question β€” "give me the sample at index i" β€” and knows nothing about batches. A DataLoader wraps a Dataset and handles everything else: grouping samples into batches, shuffling the order each epoch, and loading in parallel with background workers so your GPU never sits idle waiting for data. Layered on top, transforms clean and augment each sample as it's fetched. Get comfortable with this Dataset β†’ DataLoader β†’ training-loop pipeline and you'll be ready for the training loop itself in Module 05.

This module is split into three sub-modules β€” work through them in order.

πŸ“š Sub-Modules

# Sub-Module What you'll learn
01 Dataset Basics The Dataset interface and writing custom classes with __len__ and __getitem__
02 The DataLoader Batching, shuffling, num_workers, and iterating batches for training
03 Transforms torchvision transforms for preprocessing and data augmentation
04 Splits & Built-in Datasets Train/val/test splitting with random_split and using ready-made datasets like MNIST

🎯 By the end of this module, you'll be able to...

  • Write a custom Dataset that serves your own data one sample at a time.
  • Wrap it in a DataLoader to get shuffled, batched, parallel-loaded data.
  • Apply transforms to normalize and augment inputs before they reach the model.
  • Split data into train/validation/test sets and load ready-made datasets end-to-end.

βœ… Prerequisites

Tensors and device placement from Module 01, and a model to feed the data into from Module 03.


⬅️ Prev module: 03 Β· Neural Networks Β· ➑️ Next module: 05 Β· The Training Loop