Background
Sections
IntroductionModule 01 β€” Tensors 🧊01 Β· Creating Tensors 🧊02 Β· Indexing & Reshaping πŸ”ͺ03 Β· Tensor Math βž—04 Β· Device Placement πŸ–₯️⚑Module 02 β€” Autograd βš™οΈ01 Β· The Computational Graph πŸ•ΈοΈ02 Β· Backward Pass & Gradients ⬅️03 Β· Turning Autograd Off πŸ›‘04 Β· Gradient Gotchas πŸͺ€Module 03 β€” Neural Networks01 Β· The `nn.Module` Basics 🧱02 Β· Common Layers 🧩03 Β· Building a Network πŸ—οΈ04 Β· Inspecting Models πŸ”Module 04 β€” Data Handling πŸ—‚οΈ01 Β· Dataset Basics πŸ“‡02 Β· The DataLoader 🚚03 Β· Transforms 🎨04 Β· Splits & Built-in Datasets βœ‚οΈModule 05 β€” The Training Loop πŸ”01 Β· Loss Functions 🎯02 Β· Optimizers βš™οΈ03 Β· The Training Loop πŸ”04 Β· Evaluation & Metrics πŸ“ŠModule 06 β€” Saving & Loading πŸ’Ύ01 Β· `state_dict` Basics πŸ’Ύ02 Β· Checkpoints & Resuming ⏸️03 Β· Loading for Inference πŸš€04 Β· Best Model & Early Stopping πŸ…Module 07 β€” Computer Vision πŸ‘οΈ01 Β· Convolutions πŸ”²02 Β· CNN Architecture πŸ›οΈ03 Β· Transfer Learning πŸ”04 Β· Image Classification Project πŸ§ͺModule 08 β€” NLP & Transformers πŸ’¬01 Β· Text Data & Tokenization πŸ”€02 Β· Embeddings 🧭03 Β· Recurrent Layers & LSTMs πŸ”„04 Β· Intro to Transformers ⚑Module 09 β€” Ecosystem: PyTorch Lightning ⚑01 Β· Why Lightning? πŸ€”02 Β· The LightningModule 🧩03 Β· The Trainer πŸŽ›οΈ04 Β· DataModules & Callbacks 🧰Module 10 β€” Deployment & Optimization πŸš€01 Β· Exporting Models πŸ“¦02 Β· `torch.compile` ⚑03 Β· Inference Optimization πŸͺΆ04 Β· Serving & Next Steps πŸŽ“

Module 08 β€” NLP & Transformers πŸ’¬

3 min read

Images are grids of numbers; text is a sequence of symbols. This module teaches PyTorch to read β€” from turning words into vectors to the architecture behind modern language models.

In Module 07 you handled images, where the data arrives as neat numeric grids. Language is fundamentally different, and that difference is what makes Natural Language Processing (NLP) its own discipline. Text is a sequence of discrete symbols β€” words or sub-words β€” with no inherent numeric meaning, of variable length, where order carries enormous significance ("dog bites man" versus "man bites dog"). Before a neural network can touch it, you must convert raw strings into numbers, and you need architectures that respect sequence and context. This module walks that path: first the plumbing of getting text into tensors, then the representations and models that let a network actually understand it.

The journey mirrors the field's own evolution. You'll start by tokenizing text and looking words up in a vocabulary, then learn embeddings β€” dense, learnable vectors that capture meaning far better than crude one-hot codes. From there you'll meet recurrent networks (LSTMs), which read a sequence one step at a time while carrying a memory, the workhorse of NLP for years. Finally you'll get an introduction to the Transformer β€” the attention-based architecture that replaced recurrence and powers today's large language models. By the end you'll understand the arc from raw text to the models making headlines.

This module is split into three sub-modules β€” work through them in order.

πŸ“š Sub-Modules

# Sub-Module What you'll learn
01 Text Data & Tokenization Tokenizing text, building a vocabulary, numericalizing, and padding sequences into batches
02 Embeddings nn.Embedding and why dense learnable vectors beat one-hot encodings
03 Recurrent Layers & LSTMs RNNs, the vanishing-gradient problem, and nn.LSTM for sequences
04 Intro to Transformers The attention idea, why Transformers replaced RNNs, and nn.Transformer basics

🎯 By the end of this module, you'll be able to...

  • Convert raw text into padded tensors a model can consume.
  • Represent tokens with learnable embeddings instead of sparse one-hot vectors.
  • Build a sequence model with nn.LSTM and explain why LSTMs beat vanilla RNNs.
  • Explain self-attention and recognize the building blocks of nn.Transformer.

βœ… Prerequisites

All of Phases 1–2, plus comfort with building models (Module 03) and the data pipeline (Module 04). The collate_fn/padding idea from Module 04's DataLoader sub-module returns here.


⬅️ Prev module: 07 Β· Computer Vision Β· ➑️ Next module: 09 Β· Ecosystem: Lightning