contentintech
Advanced~20 min read

NLP

A modern tour of natural language processing from text preprocessing and TF-IDF to word embeddings, transformers, and fine-tuning models with Hugging Face.

nlptransformersembeddingshuggingface

What Is NLP?

Natural Language Processing (NLP) is the field that teaches machines to read, understand, and generate human language. Modern NLP is dominated by transformer architectures, but the core pipeline still starts with turning messy text into clean numerical representations a model can consume.

This guide moves from classical foundations (tokenization, TF-IDF) through embeddings and sequence models to the transformer stack that powers today's large language models.

Text Preprocessing

Before any model sees text, it is normalized and split into units. Key steps are tokenization, stopword removal, and stemming or lemmatization.

Tokenization

Tokenization splits a string into tokens (words, subwords, or characters). Classical NLP uses word tokenizers; modern transformers use subword tokenizers (BPE, WordPiece) so unknown words are broken into known pieces.

python
import nltk
nltk.download("punkt")
nltk.download("stopwords")
nltk.download("wordnet")

from nltk.tokenize import word_tokenize
text = "Transformers changed how machines understand language."
tokens = word_tokenize(text.lower())
# ['transformers', 'changed', 'how', 'machines', ...]

Stopwords, Stemming & Lemmatization

Stopwords (the, is, and) carry little signal for many tasks and are often removed. Stemming chops suffixes crudely (running → run), while lemmatization maps a word to its dictionary form using morphology (better → good).

python
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer, WordNetLemmatizer

stop = set(stopwords.words("english"))
tokens = [t for t in tokens if t.isalpha() and t not in stop]

stemmer = PorterStemmer()
lemmatizer = WordNetLemmatizer()
print(stemmer.stem("machines"))      # machin
print(lemmatizer.lemmatize("machines"))  # machine

Tip

Prefer lemmatization over stemming when downstream readability matters. For transformer pipelines, skip stemming entirely: the subword tokenizer and pretrained weights handle morphology far better.

Bag-of-Words & TF-IDF

Bag-of-Words (BoW) represents a document as raw word counts, ignoring order. TF-IDF reweights those counts so common words shrink and distinctive words grow.

The TF-IDF weight for term t in document d is tf(t,d) * log(N / df(t)), where N is the number of documents and df(t) is how many documents contain the term.

python
from sklearn.feature_extraction.text import TfidfVectorizer

corpus = [
    "cats chase mice",
    "dogs chase cats",
    "mice fear cats and dogs",
]
vec = TfidfVectorizer(stop_words="english")
X = vec.fit_transform(corpus)   # sparse (3 x vocab) matrix
print(vec.get_feature_names_out())
print(X.toarray().round(2))
RepresentationCaptures Order?Dense/SparseSemantics
Bag-of-WordsNoSparseNone
TF-IDFNoSparseWeak
Word2Vec / GloVeNoDenseStatic
Transformer (BERT)YesDenseContextual

Word Embeddings

Word2Vec and GloVe map each word to a dense vector such that similar words sit near each other. Word2Vec learns from local context windows (skip-gram / CBOW); GloVe factorizes a global co-occurrence matrix. The famous property: king - man + woman ≈ queen.

python
from gensim.models import Word2Vec

sentences = [["cats", "chase", "mice"], ["dogs", "chase", "cats"]]
model = Word2Vec(sentences, vector_size=100, window=5,
                 min_count=1, sg=1)   # sg=1 -> skip-gram
vec = model.wv["cats"]           # 100-dim vector
model.wv.most_similar("cats", topn=3)

The key limitation: these are static embeddings. The word "bank" gets one vector whether it means a riverbank or a financial bank. Transformers fixed this with context.

Sequence Models: RNNs & LSTMs

Before transformers, RNNs processed tokens one at a time, carrying a hidden state. LSTMs and GRUs added gating to remember long-range dependencies and fight vanishing gradients. They were state-of-the-art but slow: strictly sequential, so they cannot parallelize across a sequence.

python
import torch.nn as nn

lstm = nn.LSTM(input_size=100, hidden_size=256,
               num_layers=2, batch_first=True, bidirectional=True)
# output, (h_n, c_n) = lstm(embedded_sequence)

Transformers & Attention

The transformer replaced recurrence with self-attention, letting every token attend to every other token in parallel. Scaled dot-product attention is defined as softmax(Q Kᵀ / sqrt(d_k)) V, where Q, K, and V are learned query, key, and value projections.

Multi-head attention runs several attention computations in parallel and concatenates them, letting the model capture different relationships (syntax, coreference, topic). Positional encodings inject word order since attention is order-agnostic.

Encoder vs Decoder Models

FamilyArchitectureBest For
BERT / RoBERTaEncoder-onlyClassification, NER, embeddings
GPT / LlamaDecoder-onlyText generation, chat
T5 / BARTEncoder-decoderTranslation, summarization

BERT is trained with masked language modeling (predict hidden tokens) and reads bidirectionally. GPT-style models are trained to predict the next token left-to-right, which makes them natural generators.

Hugging Face Transformers

The transformers library is the standard toolkit. The fastest entry point is the pipeline abstraction, which bundles tokenizer + model + post-processing.

python
from transformers import pipeline

# Sentiment classification
clf = pipeline("sentiment-analysis")
clf("Hugging Face makes NLP delightful.")
# [{'label': 'POSITIVE', 'score': 0.9994}]

# Named entity recognition
ner = pipeline("ner", aggregation_strategy="simple")
ner("Ada Lovelace worked in London.")

# Summarization
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
summarizer(long_article, max_length=120, min_length=30)

Tokenizer + Model Directly

For control, load the tokenizer and model separately. The tokenizer converts text to input IDs and attention masks; the model returns logits.

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

name = "distilbert-base-uncased-finetuned-sst-2-english"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForSequenceClassification.from_pretrained(name)

inputs = tokenizer("I love transformers!", return_tensors="pt")
with torch.no_grad():
    logits = model(**inputs).logits
probs = torch.softmax(logits, dim=-1)
label = model.config.id2label[probs.argmax().item()]

Fine-Tuning (Glimpse)

Fine-tuning adapts a pretrained model to your task with the Trainer API. In 2026, parameter-efficient methods like LoRA (via the peft library) are the default for large models, updating a small set of adapter weights instead of all parameters.

python
from transformers import TrainingArguments, Trainer

args = TrainingArguments(
    output_dir="./out",
    eval_strategy="epoch",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    num_train_epochs=3,
)
trainer = Trainer(model=model, args=args,
                  train_dataset=train_ds, eval_dataset=val_ds)
trainer.train()

Rule of Thumb

Try a zero-shot or few-shot LLM prompt first. If accuracy or latency is insufficient, fine-tune a smaller encoder model with LoRA. Only full fine-tune when you have plenty of labeled data and compute.

Common NLP Tasks

TaskPipeline NameTypical Model
Text classificationtext-classificationDistilBERT, RoBERTa
Named entity recognitionnerBERT-NER, spaCy
Question answeringquestion-answeringDeBERTa, RoBERTa
SummarizationsummarizationBART, T5
Text generationtext-generationGPT, Llama, Mistral

Practice Exercises

  1. Build a preprocessing function that lowercases text, removes stopwords and punctuation, and lemmatizes tokens. Compare vocabulary size before and after.
  2. Fit a TfidfVectorizer plus logistic regression on a movie-review dataset and report accuracy. Which words have the highest TF-IDF weights per class?
  3. Train a small Word2Vec model on any text corpus and verify at least one analogy (e.g. most_similar(positive=["king","woman"], negative=["man"])).
  4. Use a Hugging Face pipeline to run sentiment analysis and NER on 5 news headlines. Inspect the confidence scores.
  5. Load a BERT tokenizer and print the subword tokens for the word "unbelievability". Explain how the pieces were formed.
  6. Fine-tune (or LoRA-tune) DistilBERT on a two-class text dataset for 3 epochs and compare against your TF-IDF baseline.

Section navigation