contentintech
Learn/data science/Deep Learning
Advanced~20 min read

Deep Learning

Neural network fundamentals through CNNs, RNNs, and transformers, with a full PyTorch training loop.

pytorchneural-networkscnntransformers

What Makes Learning "Deep"

Deep learning stacks many layers of simple, differentiable units so that a network learns a hierarchy of representations directly from raw data — pixels, tokens, waveforms — instead of hand-crafted features. The whole system is trained end-to-end with gradient descent and backpropagation. This guide uses PyTorch.

The Neuron and Layers

A single artificial neuron computes a weighted sum plus bias, then applies a nonlinear activation: a = f(w·x + b). Stacking neurons into layers, and layers into a network, lets the model approximate arbitrarily complex functions. A fully-connected (dense) layer is just a matrix multiply followed by an activation.

python
import torch
import torch.nn as nn

layer = nn.Linear(in_features=128, out_features=64)  # W: 64x128, b: 64
x = torch.randn(32, 128)   # batch of 32
h = torch.relu(layer(x))   # -> (32, 64)

Activation Functions

Without a nonlinear activation, stacking linear layers collapses to a single linear map. Activations give the network its expressive power.

ActivationRangeUse
ReLU[0, ∞)Default hidden layers
GELU(−, ∞)Transformers
Sigmoid(0, 1)Binary output prob
Tanh(−1, 1)RNN gates
Softmax(0, 1), Σ=1Multi-class output

Forward & Backpropagation

The forward pass pushes inputs through the layers to produce a prediction and a scalar loss. Backpropagation applies the chain rule backward through the computation graph to compute the gradient of the loss with respect to every parameter. In PyTorch, autograd builds this graph automatically; loss.backward() fills every parameter's .grad.

python
w = torch.tensor([2.0], requires_grad=True)
x = torch.tensor([3.0])
y = (w * x).sum()      # forward
y.backward()           # backprop -> dw = x
print(w.grad)          # tensor([3.])

Loss Functions

The loss quantifies how wrong a prediction is. Regression typically uses mean squared error; classification uses cross-entropy.

python
mse = nn.MSELoss()                 # regression
bce = nn.BCEWithLogitsLoss()       # binary (logits in)
ce  = nn.CrossEntropyLoss()        # multi-class (logits + int labels)

Numerical stability

Prefer BCEWithLogitsLoss / CrossEntropyLoss over applying sigmoid/softmax then a separate loss. The fused versions use the log-sum-exp trick and are far more numerically stable.

Gradient Descent & Optimizers

Optimizers use gradients to update parameters: θ ← θ − η · ∇L. Plain SGD is simple; momentum and adaptive methods like Adam converge faster and more reliably.

OptimizerIdeaNotes
SGDStep down the gradientAdd momentum=0.9
AdamAdaptive per-param LRGreat default
AdamWAdam + decoupled decayStandard for transformers
RMSpropNormalize by grad RMSCommon for RNNs

Convolutional Neural Networks (CNNs)

CNNs exploit spatial structure with weight-sharing convolutional filters. Each filter slides over the input detecting local patterns (edges, then textures, then objects). Pooling downsamples, and stacked conv blocks build a feature hierarchy — ideal for images.

python
cnn = nn.Sequential(
    nn.Conv2d(3, 32, kernel_size=3, padding=1),
    nn.BatchNorm2d(32),
    nn.ReLU(),
    nn.MaxPool2d(2),                 # 32x16x16
    nn.Conv2d(32, 64, 3, padding=1),
    nn.ReLU(),
    nn.AdaptiveAvgPool2d(1),         # 64x1x1
    nn.Flatten(),
    nn.Linear(64, 10),
)

RNNs and LSTMs

Recurrent networks process sequences one step at a time, maintaining a hidden state that carries context. Vanilla RNNs struggle with long-range dependencies due to vanishing gradients; LSTMs add gated memory cells (input, forget, output gates) to preserve information over long spans.

python
lstm = nn.LSTM(input_size=64, hidden_size=128,
               num_layers=2, batch_first=True, dropout=0.2)
seq = torch.randn(32, 50, 64)        # (batch, time, feat)
out, (h_n, c_n) = lstm(seq)          # out: (32, 50, 128)

A Glimpse of Transformers

Transformers replace recurrence with self-attention: every token attends to every other token in parallel, weighting them by learned relevance. Attention is softmax(QKᵀ / √d) V. This parallelism and long-range modeling underpin modern LLMs and vision transformers.

python
attn = nn.MultiheadAttention(embed_dim=256, num_heads=8, batch_first=True)
x = torch.randn(32, 20, 256)          # (batch, seq, dim)
out, weights = attn(x, x, x)          # self-attention

# A full encoder block in one line:
block = nn.TransformerEncoderLayer(d_model=256, nhead=8, batch_first=True)

Regularization: Dropout & BatchNorm

Dropout randomly zeros a fraction of activations during training, preventing co-adaptation. Batch normalization standardizes layer inputs per mini-batch, stabilizing and speeding up training. Both behave differently in train vs eval mode — always toggle with model.train() / model.eval().

A Full PyTorch Training Loop

Everything comes together in the canonical loop: zero gradients, forward, compute loss, backward, step. Wrap validation in torch.no_grad().

python
import torch
import torch.nn as nn
from torch.utils.data import DataLoader

device = "cuda" if torch.cuda.is_available() else "cpu"

model = nn.Sequential(
    nn.Linear(784, 256), nn.ReLU(), nn.Dropout(0.3),
    nn.Linear(256, 10),
).to(device)

criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=1e-2)
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=10)

def train_epoch(loader):
    model.train()
    total = 0.0
    for xb, yb in loader:
        xb, yb = xb.to(device), yb.to(device)
        optimizer.zero_grad()
        logits = model(xb)
        loss = criterion(logits, yb)
        loss.backward()
        nn.utils.clip_grad_norm_(model.parameters(), 1.0)  # stability
        optimizer.step()
        total += loss.item() * xb.size(0)
    scheduler.step()
    return total / len(loader.dataset)

@torch.no_grad()
def evaluate(loader):
    model.eval()
    correct = 0
    for xb, yb in loader:
        xb, yb = xb.to(device), yb.to(device)
        preds = model(xb).argmax(dim=1)
        correct += (preds == yb).sum().item()
    return correct / len(loader.dataset)

for epoch in range(10):
    train_loss = train_epoch(train_loader)
    acc = evaluate(val_loader)
    print(f"epoch {epoch}: loss={train_loss:.4f} val_acc={acc:.4f}")

torch.save(model.state_dict(), "model.pt")

Do not forget zero_grad

PyTorch accumulates gradients by default. Call optimizer.zero_grad() every step or gradients from previous batches will pile up and corrupt training.

Practice Exercises

  1. Build a 2-layer MLP for MNIST and train it with the loop above; report validation accuracy.
  2. Swap CrossEntropyLoss for a manual softmax + NLL and explain why the fused version is more stable.
  3. Add dropout and batch norm to a CNN and measure the effect on the train/val accuracy gap.
  4. Compare SGD with momentum against AdamW on the same model and plot the loss curves.
  5. Implement scaled dot-product attention from scratch and verify it matches nn.MultiheadAttention.
  6. Train an LSTM on a character-level sequence task and generate a short continuation.
  7. Add gradient clipping and a cosine learning-rate schedule, then observe their effect on convergence.

Section navigation