contentintech
Intermediate~20 min read

MLOps

A practical guide to the ML lifecycle covering experiment tracking, versioning, pipelines, serving, CI/CD, monitoring, and drift detection.

mlopsmlflowdockerdeployment

What Is MLOps?

MLOps (Machine Learning Operations) applies DevOps principles to machine learning: reproducibility, automation, testing, and monitoring across the full model lifecycle. The goal is to move models from a notebook to reliable production and keep them healthy over time.

Unlike traditional software, ML systems have three moving parts that can each break: code, data, and model. MLOps versions and monitors all three.

The ML Lifecycle

A production ML system flows through repeatable stages, and each loops back as data and requirements change.

StageGoalTools
Data ingestion & validationClean, versioned dataDVC, Great Expectations
ExperimentationTrack runs & metricsMLflow, Weights & Biases
Training pipelineReproducible retrainingAirflow, Prefect, Kubeflow
Packaging & servingExpose predictionsFastAPI, Docker, BentoML
MonitoringDetect drift & decayEvidently, Prometheus

Experiment Tracking with MLflow

Every training run should log its parameters, metrics, and artifacts so you can compare and reproduce. MLflow is the de facto open-source tracker.

python
import mlflow
from sklearn.ensemble import RandomForestClassifier

mlflow.set_experiment("churn-prediction")

with mlflow.start_run(run_name="rf-baseline"):
    params = {"n_estimators": 200, "max_depth": 8}
    mlflow.log_params(params)

    model = RandomForestClassifier(**params).fit(X_train, y_train)
    acc = model.score(X_test, y_test)

    mlflow.log_metric("accuracy", acc)
    mlflow.sklearn.log_model(model, artifact_path="model")

Launch the UI with mlflow ui to compare runs side by side. Autologging (mlflow.autolog()) captures params and metrics automatically for popular frameworks.

Data & Model Versioning with DVC

Git handles code, but datasets and model files are too large. DVC (Data Version Control) stores lightweight pointers in Git while pushing the actual bytes to remote storage (S3, GCS).

bash
dvc init
dvc add data/train.csv          # creates data/train.csv.dvc
git add data/train.csv.dvc .gitignore
git commit -m "Track training data"

dvc remote add -d storage s3://my-bucket/dvcstore
dvc push                        # upload data to remote
dvc pull                        # teammate reproduces exact data

Reproducibility

A model is only reproducible if you can pin the code commit, the data version, and the environment (pinned dependencies). DVC ties the data hash to the Git commit so git checkout + dvc checkout restores both.

Reproducible Pipelines

DVC pipelines declare stages with inputs and outputs in dvc.yaml. DVC only reruns a stage when its dependencies change, giving cached, deterministic runs.

yaml
stages:
  prepare:
    cmd: python src/prepare.py
    deps: [data/raw.csv, src/prepare.py]
    outs: [data/clean.csv]
  train:
    cmd: python src/train.py
    deps: [data/clean.csv, src/train.py]
    params: [train.n_estimators, train.max_depth]
    outs: [models/model.pkl]
    metrics: [metrics.json]

Run the whole DAG with dvc repro and compare experiments with dvc exp show.

Model Packaging & Serving

Wrap the model in a REST API so applications can request predictions over HTTP. FastAPI is the standard choice: fast, async, with automatic validation and OpenAPI docs.

python
from fastapi import FastAPI
from pydantic import BaseModel
import joblib

app = FastAPI(title="Churn API")
model = joblib.load("models/model.pkl")

class Features(BaseModel):
    tenure: float
    monthly_charges: float
    contract_type: int

@app.post("/predict")
def predict(f: Features):
    x = [[f.tenure, f.monthly_charges, f.contract_type]]
    proba = float(model.predict_proba(x)[0][1])
    return {"churn_probability": proba,
            "prediction": int(proba > 0.5)}

@app.get("/health")
def health():
    return {"status": "ok"}

Run locally with uvicorn app:app --reload. Now package it so it runs identically anywhere.

Dockerizing the Service

dockerfile
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
bash
docker build -t churn-api:1.0 .
docker run -p 8000:8000 churn-api:1.0

CI/CD for ML

Continuous integration for ML runs data validation, unit tests, and (optionally) a training + evaluation gate on every push. A model is only promoted if it beats the current production baseline.

yaml
# .github/workflows/ci.yml
name: ml-ci
on: [push]
jobs:
  test-and-train:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -r requirements.txt
      - run: pytest tests/
      - run: dvc pull && dvc repro
      - run: python src/evaluate.py --min-accuracy 0.85

Model Registry

A model registry is the source of truth for which model version is in which stage (Staging, Production, Archived). MLflow's registry manages this transition lifecycle.

python
from mlflow import MlflowClient

client = MlflowClient()
mv = mlflow.register_model("runs:/<run_id>/model", "churn-model")
client.set_registered_model_alias(
    "churn-model", alias="production", version=mv.version)

# Serving code loads by alias, not a hardcoded version
model = mlflow.pyfunc.load_model("models:/churn-model@production")

Monitoring & Drift Detection

Models degrade silently as the world changes. Two failure modes to watch: data drift (input distribution shifts) and concept drift (the relationship between inputs and target changes).

python
from evidently import Report
from evidently.presets import DataDriftPreset

report = Report(metrics=[DataDriftPreset()])
result = report.run(reference_data=train_df, current_data=live_df)
result.save_html("drift_report.html")

Track prediction latency, error rates, and feature statistics with Prometheus + Grafana. Trigger retraining when drift crosses a threshold or accuracy drops below the SLA.

Feature Stores

A feature store (e.g. Feast) centralizes feature computation so training and serving use identical logic, eliminating training-serving skew. It offers an offline store for batch training and an online store for low-latency inference.

Orchestration

Pipelines need scheduling, retries, and dependency management. Airflow and Prefect are the common orchestrators.

python
from prefect import flow, task

@task(retries=2)
def extract(): ...
@task
def train(data): ...

@flow(name="daily-retrain")
def pipeline():
    data = extract()
    train(data)

if __name__ == "__main__":
    pipeline.serve(cron="0 2 * * *")   # 2am daily

Maturity Ladder

Level 0 is manual notebooks. Level 1 automates the training pipeline. Level 2 adds full CI/CD with automated retraining, testing, and deployment. Advance one rung at a time rather than jumping to full automation prematurely.

Practice Exercises

  1. Wrap a scikit-learn training script in an MLflow run that logs at least three parameters, one metric, and the model artifact. Open the MLflow UI and compare two runs.
  2. Initialize DVC in a repo, track a CSV dataset, push it to a local remote, then simulate a teammate with dvc pull.
  3. Build a FastAPI service with /predict and /health endpoints, then containerize it with a Dockerfile and run the image.
  4. Write a dvc.yaml with prepare and train stages and confirm dvc repro caches unchanged stages.
  5. Register a model in the MLflow registry, promote a version to the production alias, and load it by alias.
  6. Generate an Evidently data-drift report comparing a reference set to a shifted current set. Which features drift most?

Section navigation