What Makes Learning "Deep"
Deep learning stacks many layers of simple, differentiable units so that a network learns a hierarchy of representations directly from raw data — pixels, tokens, waveforms — instead of hand-crafted features. The whole system is trained end-to-end with gradient descent and backpropagation. This guide uses PyTorch.
The Neuron and Layers
A single artificial neuron computes a weighted sum plus bias, then applies a nonlinear activation: a = f(w·x + b). Stacking neurons into layers, and layers into a network, lets the model approximate arbitrarily complex functions. A fully-connected (dense) layer is just a matrix multiply followed by an activation.
import torch
import torch.nn as nn
layer = nn.Linear(in_features=128, out_features=64) # W: 64x128, b: 64
x = torch.randn(32, 128) # batch of 32
h = torch.relu(layer(x)) # -> (32, 64)
Activation Functions
Without a nonlinear activation, stacking linear layers collapses to a single linear map. Activations give the network its expressive power.
| Activation | Range | Use |
|---|---|---|
| ReLU | [0, ∞) | Default hidden layers |
| GELU | (−, ∞) | Transformers |
| Sigmoid | (0, 1) | Binary output prob |
| Tanh | (−1, 1) | RNN gates |
| Softmax | (0, 1), Σ=1 | Multi-class output |
Forward & Backpropagation
The forward pass pushes inputs through the layers to produce a prediction and a scalar loss. Backpropagation applies the chain rule backward through the computation graph to compute the gradient of the loss with respect to every parameter. In PyTorch, autograd builds this graph automatically; loss.backward() fills every parameter's .grad.
w = torch.tensor([2.0], requires_grad=True)
x = torch.tensor([3.0])
y = (w * x).sum() # forward
y.backward() # backprop -> dw = x
print(w.grad) # tensor([3.])
Loss Functions
The loss quantifies how wrong a prediction is. Regression typically uses mean squared error; classification uses cross-entropy.
mse = nn.MSELoss() # regression
bce = nn.BCEWithLogitsLoss() # binary (logits in)
ce = nn.CrossEntropyLoss() # multi-class (logits + int labels)
Numerical stability
Prefer BCEWithLogitsLoss / CrossEntropyLoss over applying sigmoid/softmax then a separate loss. The fused versions use the log-sum-exp trick and are far more numerically stable.
Gradient Descent & Optimizers
Optimizers use gradients to update parameters: θ ← θ − η · ∇L. Plain SGD is simple; momentum and adaptive methods like Adam converge faster and more reliably.
| Optimizer | Idea | Notes |
|---|---|---|
| SGD | Step down the gradient | Add momentum=0.9 |
| Adam | Adaptive per-param LR | Great default |
| AdamW | Adam + decoupled decay | Standard for transformers |
| RMSprop | Normalize by grad RMS | Common for RNNs |
Convolutional Neural Networks (CNNs)
CNNs exploit spatial structure with weight-sharing convolutional filters. Each filter slides over the input detecting local patterns (edges, then textures, then objects). Pooling downsamples, and stacked conv blocks build a feature hierarchy — ideal for images.
cnn = nn.Sequential(
nn.Conv2d(3, 32, kernel_size=3, padding=1),
nn.BatchNorm2d(32),
nn.ReLU(),
nn.MaxPool2d(2), # 32x16x16
nn.Conv2d(32, 64, 3, padding=1),
nn.ReLU(),
nn.AdaptiveAvgPool2d(1), # 64x1x1
nn.Flatten(),
nn.Linear(64, 10),
)
RNNs and LSTMs
Recurrent networks process sequences one step at a time, maintaining a hidden state that carries context. Vanilla RNNs struggle with long-range dependencies due to vanishing gradients; LSTMs add gated memory cells (input, forget, output gates) to preserve information over long spans.
lstm = nn.LSTM(input_size=64, hidden_size=128,
num_layers=2, batch_first=True, dropout=0.2)
seq = torch.randn(32, 50, 64) # (batch, time, feat)
out, (h_n, c_n) = lstm(seq) # out: (32, 50, 128)
A Glimpse of Transformers
Transformers replace recurrence with self-attention: every token attends to every other token in parallel, weighting them by learned relevance. Attention is softmax(QKᵀ / √d) V. This parallelism and long-range modeling underpin modern LLMs and vision transformers.
attn = nn.MultiheadAttention(embed_dim=256, num_heads=8, batch_first=True)
x = torch.randn(32, 20, 256) # (batch, seq, dim)
out, weights = attn(x, x, x) # self-attention
# A full encoder block in one line:
block = nn.TransformerEncoderLayer(d_model=256, nhead=8, batch_first=True)
Regularization: Dropout & BatchNorm
Dropout randomly zeros a fraction of activations during training, preventing co-adaptation. Batch normalization standardizes layer inputs per mini-batch, stabilizing and speeding up training. Both behave differently in train vs eval mode — always toggle with model.train() / model.eval().
A Full PyTorch Training Loop
Everything comes together in the canonical loop: zero gradients, forward, compute loss, backward, step. Wrap validation in torch.no_grad().
import torch
import torch.nn as nn
from torch.utils.data import DataLoader
device = "cuda" if torch.cuda.is_available() else "cpu"
model = nn.Sequential(
nn.Linear(784, 256), nn.ReLU(), nn.Dropout(0.3),
nn.Linear(256, 10),
).to(device)
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=1e-2)
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=10)
def train_epoch(loader):
model.train()
total = 0.0
for xb, yb in loader:
xb, yb = xb.to(device), yb.to(device)
optimizer.zero_grad()
logits = model(xb)
loss = criterion(logits, yb)
loss.backward()
nn.utils.clip_grad_norm_(model.parameters(), 1.0) # stability
optimizer.step()
total += loss.item() * xb.size(0)
scheduler.step()
return total / len(loader.dataset)
@torch.no_grad()
def evaluate(loader):
model.eval()
correct = 0
for xb, yb in loader:
xb, yb = xb.to(device), yb.to(device)
preds = model(xb).argmax(dim=1)
correct += (preds == yb).sum().item()
return correct / len(loader.dataset)
for epoch in range(10):
train_loss = train_epoch(train_loader)
acc = evaluate(val_loader)
print(f"epoch {epoch}: loss={train_loss:.4f} val_acc={acc:.4f}")
torch.save(model.state_dict(), "model.pt")
Do not forget zero_grad
PyTorch accumulates gradients by default. Call optimizer.zero_grad() every step or gradients from previous batches will pile up and corrupt training.
Practice Exercises
- Build a 2-layer MLP for MNIST and train it with the loop above; report validation accuracy.
- Swap CrossEntropyLoss for a manual softmax + NLL and explain why the fused version is more stable.
- Add dropout and batch norm to a CNN and measure the effect on the train/val accuracy gap.
- Compare SGD with momentum against AdamW on the same model and plot the loss curves.
- Implement scaled dot-product attention from scratch and verify it matches nn.MultiheadAttention.
- Train an LSTM on a character-level sequence task and generate a short continuation.
- Add gradient clipping and a cosine learning-rate schedule, then observe their effect on convergence.