The Machine Learning Workflow
Machine learning is the practice of fitting models to data so they generalize to unseen examples. A production workflow almost always follows the same loop: frame the problem, collect and clean data, split it, engineer features, train, evaluate, tune, and deploy. Everything below fits into that loop.
We use scikit-learn for classical ML because its estimator API — .fit(), .predict(), .transform() — is consistent across every algorithm.
Supervised vs Unsupervised Learning
Supervised learning uses labeled data: each example has a known target. Regression predicts a continuous value; classification predicts a category. Unsupervised learning finds structure in unlabeled data — clustering, dimensionality reduction, anomaly detection.
| Paradigm | Target | Examples |
|---|---|---|
| Supervised (regression) | Continuous | Linear regression, gradient boosting |
| Supervised (classification) | Discrete label | Logistic regression, random forest |
| Unsupervised | None | k-means, PCA |
Train/Test Split
Never evaluate on data the model trained on. Hold out a test set to estimate generalization. Use stratify to preserve class balance in classification.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
Linear Regression
Linear regression fits a line (or hyperplane) minimizing squared error. The prediction is ŷ = w·x + b, and training minimizes the mean squared error loss MSE = (1/n) Σ (yᵢ − ŷᵢ)².
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score
model = LinearRegression().fit(X_train, y_train)
pred = model.predict(X_test)
print(model.coef_, model.intercept_)
print("R2:", r2_score(y_test, pred))
print("RMSE:", mean_squared_error(y_test, pred, squared=False))
Logistic Regression
Despite the name, logistic regression is a classifier. It passes a linear score through the sigmoid σ(z) = 1 / (1 + e⁻ᶻ) to output a probability, and minimizes log loss (binary cross-entropy) −Σ [y·log(p) + (1−y)·log(1−p)].
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(C=1.0, max_iter=1000).fit(X_train, y_train)
proba = clf.predict_proba(X_test)[:, 1] # P(class = 1)
labels = clf.predict(X_test)
Decision Trees & Random Forests
A decision tree recursively splits the feature space to reduce impurity (Gini or entropy). Single trees overfit easily; a random forest averages many decorrelated trees (bagging + feature subsampling) to reduce variance.
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(
n_estimators=300,
max_depth=None,
min_samples_leaf=2,
max_features="sqrt",
n_jobs=-1,
random_state=42,
).fit(X_train, y_train)
importances = rf.feature_importances_
Gradient Boosting
Boosting builds trees sequentially, each correcting the errors of the ensemble so far by fitting the negative gradient of the loss. Modern libraries like XGBoost and LightGBM dominate tabular competitions; scikit-learn ships HistGradientBoosting.
from sklearn.ensemble import HistGradientBoostingClassifier
gb = HistGradientBoostingClassifier(
learning_rate=0.05,
max_iter=500,
max_leaf_nodes=31,
early_stopping=True,
validation_fraction=0.1,
).fit(X_train, y_train)
Bagging vs Boosting
Bagging (random forest) trains trees in parallel to reduce variance. Boosting trains trees sequentially to reduce bias. Boosting usually wins on accuracy but is more sensitive to hyperparameters and noise.
k-Nearest Neighbors
kNN is a lazy, instance-based method: to classify a point, look at its k nearest neighbors and vote. It has no training phase but needs feature scaling and can be slow at prediction time.
from sklearn.neighbors import KNeighborsClassifier
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
knn = make_pipeline(
StandardScaler(),
KNeighborsClassifier(n_neighbors=7, weights="distance"),
).fit(X_train, y_train)
k-Means Clustering
k-means partitions data into k clusters by iteratively assigning points to the nearest centroid and recomputing centroids to minimize within-cluster variance (inertia). Choose k with the elbow method or silhouette score.
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
km = KMeans(n_clusters=4, n_init="auto", random_state=42).fit(X)
print("inertia:", km.inertia_)
print("silhouette:", silhouette_score(X, km.labels_))
Overfitting & the Bias-Variance Tradeoff
A model that memorizes training noise overfits (low bias, high variance); one too simple to capture the signal underfits (high bias, low variance). Generalization error decomposes into bias² + variance + irreducible noise. The goal is the sweet spot between them.
Regularization (L1 & L2)
Regularization penalizes large weights to fight overfitting. L2 (Ridge) adds λ Σ wⱼ² and shrinks weights smoothly. L1 (Lasso) adds λ Σ |wⱼ| and drives some weights to exactly zero, performing feature selection.
from sklearn.linear_model import Ridge, Lasso
ridge = Ridge(alpha=1.0).fit(X_train, y_train) # L2
lasso = Lasso(alpha=0.1).fit(X_train, y_train) # L1 -> sparse coefs
Cross-Validation
A single split is noisy. k-fold cross-validation rotates the validation fold k times and averages the scores, giving a more robust estimate. Pair it with GridSearchCV to tune hyperparameters.
from sklearn.model_selection import cross_val_score, GridSearchCV
scores = cross_val_score(rf, X, y, cv=5, scoring="f1")
print(scores.mean(), scores.std())
grid = GridSearchCV(
RandomForestClassifier(random_state=42),
param_grid={"max_depth": [5, 10, None], "n_estimators": [100, 300]},
cv=5, scoring="roc_auc", n_jobs=-1,
).fit(X_train, y_train)
print(grid.best_params_, grid.best_score_)
Evaluation Metrics
Pick metrics that match the problem. Accuracy is misleading on imbalanced data — use precision, recall, F1, and ROC-AUC instead.
| Metric | Task | Meaning |
|---|---|---|
| Accuracy | Classification | Correct / total |
| Precision | Classification | TP / (TP + FP) |
| Recall | Classification | TP / (TP + FN) |
| F1 | Classification | Harmonic mean of P & R |
| ROC-AUC | Classification | Ranking quality, 0.5–1.0 |
| RMSE / MAE | Regression | Average error magnitude |
| R² | Regression | Variance explained |
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report,
)
print(classification_report(y_test, labels))
print("ROC-AUC:", roc_auc_score(y_test, proba))
Precision vs Recall
Optimize precision when false positives are costly (spam filters). Optimize recall when false negatives are costly (disease screening). F1 balances both.
Practice Exercises
- Split a dataset with an 80/20 stratified split, then train and evaluate a logistic regression model.
- Compare a single decision tree against a random forest on the same data and explain the accuracy difference in terms of bias and variance.
- Tune the
alphaof Ridge and Lasso, and report which features Lasso zeroed out. - Use 5-fold cross-validation with GridSearchCV to tune a HistGradientBoostingClassifier by ROC-AUC.
- On an imbalanced dataset, show why accuracy is misleading and report precision, recall, and F1 instead.
- Cluster an unlabeled dataset with k-means and use the silhouette score to choose the number of clusters.
- Build a scikit-learn Pipeline that scales features and fits kNN, then evaluate it with cross-validation.