contentintech
Learn/data science/Machine Learning
Intermediate~20 min read

Machine Learning

Core machine learning concepts and algorithms with scikit-learn, from regression and trees to cross-validation and metrics.

scikit-learnregressionclassificationmodel-evaluation

The Machine Learning Workflow

Machine learning is the practice of fitting models to data so they generalize to unseen examples. A production workflow almost always follows the same loop: frame the problem, collect and clean data, split it, engineer features, train, evaluate, tune, and deploy. Everything below fits into that loop.

We use scikit-learn for classical ML because its estimator API — .fit(), .predict(), .transform() — is consistent across every algorithm.

Supervised vs Unsupervised Learning

Supervised learning uses labeled data: each example has a known target. Regression predicts a continuous value; classification predicts a category. Unsupervised learning finds structure in unlabeled data — clustering, dimensionality reduction, anomaly detection.

ParadigmTargetExamples
Supervised (regression)ContinuousLinear regression, gradient boosting
Supervised (classification)Discrete labelLogistic regression, random forest
UnsupervisedNonek-means, PCA

Train/Test Split

Never evaluate on data the model trained on. Hold out a test set to estimate generalization. Use stratify to preserve class balance in classification.

python
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

Linear Regression

Linear regression fits a line (or hyperplane) minimizing squared error. The prediction is ŷ = w·x + b, and training minimizes the mean squared error loss MSE = (1/n) Σ (yᵢ − ŷᵢ)².

python
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score

model = LinearRegression().fit(X_train, y_train)
pred = model.predict(X_test)

print(model.coef_, model.intercept_)
print("R2:", r2_score(y_test, pred))
print("RMSE:", mean_squared_error(y_test, pred, squared=False))

Logistic Regression

Despite the name, logistic regression is a classifier. It passes a linear score through the sigmoid σ(z) = 1 / (1 + e⁻ᶻ) to output a probability, and minimizes log loss (binary cross-entropy) −Σ [y·log(p) + (1−y)·log(1−p)].

python
from sklearn.linear_model import LogisticRegression

clf = LogisticRegression(C=1.0, max_iter=1000).fit(X_train, y_train)
proba = clf.predict_proba(X_test)[:, 1]   # P(class = 1)
labels = clf.predict(X_test)

Decision Trees & Random Forests

A decision tree recursively splits the feature space to reduce impurity (Gini or entropy). Single trees overfit easily; a random forest averages many decorrelated trees (bagging + feature subsampling) to reduce variance.

python
from sklearn.ensemble import RandomForestClassifier

rf = RandomForestClassifier(
    n_estimators=300,
    max_depth=None,
    min_samples_leaf=2,
    max_features="sqrt",
    n_jobs=-1,
    random_state=42,
).fit(X_train, y_train)

importances = rf.feature_importances_

Gradient Boosting

Boosting builds trees sequentially, each correcting the errors of the ensemble so far by fitting the negative gradient of the loss. Modern libraries like XGBoost and LightGBM dominate tabular competitions; scikit-learn ships HistGradientBoosting.

python
from sklearn.ensemble import HistGradientBoostingClassifier

gb = HistGradientBoostingClassifier(
    learning_rate=0.05,
    max_iter=500,
    max_leaf_nodes=31,
    early_stopping=True,
    validation_fraction=0.1,
).fit(X_train, y_train)

Bagging vs Boosting

Bagging (random forest) trains trees in parallel to reduce variance. Boosting trains trees sequentially to reduce bias. Boosting usually wins on accuracy but is more sensitive to hyperparameters and noise.

k-Nearest Neighbors

kNN is a lazy, instance-based method: to classify a point, look at its k nearest neighbors and vote. It has no training phase but needs feature scaling and can be slow at prediction time.

python
from sklearn.neighbors import KNeighborsClassifier
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

knn = make_pipeline(
    StandardScaler(),
    KNeighborsClassifier(n_neighbors=7, weights="distance"),
).fit(X_train, y_train)

k-Means Clustering

k-means partitions data into k clusters by iteratively assigning points to the nearest centroid and recomputing centroids to minimize within-cluster variance (inertia). Choose k with the elbow method or silhouette score.

python
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

km = KMeans(n_clusters=4, n_init="auto", random_state=42).fit(X)
print("inertia:", km.inertia_)
print("silhouette:", silhouette_score(X, km.labels_))

Overfitting & the Bias-Variance Tradeoff

A model that memorizes training noise overfits (low bias, high variance); one too simple to capture the signal underfits (high bias, low variance). Generalization error decomposes into bias² + variance + irreducible noise. The goal is the sweet spot between them.

Regularization (L1 & L2)

Regularization penalizes large weights to fight overfitting. L2 (Ridge) adds λ Σ wⱼ² and shrinks weights smoothly. L1 (Lasso) adds λ Σ |wⱼ| and drives some weights to exactly zero, performing feature selection.

python
from sklearn.linear_model import Ridge, Lasso

ridge = Ridge(alpha=1.0).fit(X_train, y_train)   # L2
lasso = Lasso(alpha=0.1).fit(X_train, y_train)    # L1 -> sparse coefs

Cross-Validation

A single split is noisy. k-fold cross-validation rotates the validation fold k times and averages the scores, giving a more robust estimate. Pair it with GridSearchCV to tune hyperparameters.

python
from sklearn.model_selection import cross_val_score, GridSearchCV

scores = cross_val_score(rf, X, y, cv=5, scoring="f1")
print(scores.mean(), scores.std())

grid = GridSearchCV(
    RandomForestClassifier(random_state=42),
    param_grid={"max_depth": [5, 10, None], "n_estimators": [100, 300]},
    cv=5, scoring="roc_auc", n_jobs=-1,
).fit(X_train, y_train)
print(grid.best_params_, grid.best_score_)

Evaluation Metrics

Pick metrics that match the problem. Accuracy is misleading on imbalanced data — use precision, recall, F1, and ROC-AUC instead.

MetricTaskMeaning
AccuracyClassificationCorrect / total
PrecisionClassificationTP / (TP + FP)
RecallClassificationTP / (TP + FN)
F1ClassificationHarmonic mean of P & R
ROC-AUCClassificationRanking quality, 0.5–1.0
RMSE / MAERegressionAverage error magnitude
R²RegressionVariance explained
python
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report,
)

print(classification_report(y_test, labels))
print("ROC-AUC:", roc_auc_score(y_test, proba))

Precision vs Recall

Optimize precision when false positives are costly (spam filters). Optimize recall when false negatives are costly (disease screening). F1 balances both.

Practice Exercises

  1. Split a dataset with an 80/20 stratified split, then train and evaluate a logistic regression model.
  2. Compare a single decision tree against a random forest on the same data and explain the accuracy difference in terms of bias and variance.
  3. Tune the alpha of Ridge and Lasso, and report which features Lasso zeroed out.
  4. Use 5-fold cross-validation with GridSearchCV to tune a HistGradientBoostingClassifier by ROC-AUC.
  5. On an imbalanced dataset, show why accuracy is misleading and report precision, recall, and F1 instead.
  6. Cluster an unlabeled dataset with k-means and use the silhouette score to choose the number of clusters.
  7. Build a scikit-learn Pipeline that scales features and fits kNN, then evaluate it with cross-validation.

Section navigation