What Is MLOps?
MLOps (Machine Learning Operations) applies DevOps principles to machine learning: reproducibility, automation, testing, and monitoring across the full model lifecycle. The goal is to move models from a notebook to reliable production and keep them healthy over time.
Unlike traditional software, ML systems have three moving parts that can each break: code, data, and model. MLOps versions and monitors all three.
The ML Lifecycle
A production ML system flows through repeatable stages, and each loops back as data and requirements change.
| Stage | Goal | Tools |
|---|---|---|
| Data ingestion & validation | Clean, versioned data | DVC, Great Expectations |
| Experimentation | Track runs & metrics | MLflow, Weights & Biases |
| Training pipeline | Reproducible retraining | Airflow, Prefect, Kubeflow |
| Packaging & serving | Expose predictions | FastAPI, Docker, BentoML |
| Monitoring | Detect drift & decay | Evidently, Prometheus |
Experiment Tracking with MLflow
Every training run should log its parameters, metrics, and artifacts so you can compare and reproduce. MLflow is the de facto open-source tracker.
import mlflow
from sklearn.ensemble import RandomForestClassifier
mlflow.set_experiment("churn-prediction")
with mlflow.start_run(run_name="rf-baseline"):
params = {"n_estimators": 200, "max_depth": 8}
mlflow.log_params(params)
model = RandomForestClassifier(**params).fit(X_train, y_train)
acc = model.score(X_test, y_test)
mlflow.log_metric("accuracy", acc)
mlflow.sklearn.log_model(model, artifact_path="model")
Launch the UI with mlflow ui to compare runs side by side. Autologging (mlflow.autolog()) captures params and metrics automatically for popular frameworks.
Data & Model Versioning with DVC
Git handles code, but datasets and model files are too large. DVC (Data Version Control) stores lightweight pointers in Git while pushing the actual bytes to remote storage (S3, GCS).
dvc init
dvc add data/train.csv # creates data/train.csv.dvc
git add data/train.csv.dvc .gitignore
git commit -m "Track training data"
dvc remote add -d storage s3://my-bucket/dvcstore
dvc push # upload data to remote
dvc pull # teammate reproduces exact data
Reproducibility
A model is only reproducible if you can pin the code commit, the data version, and the environment (pinned dependencies). DVC ties the data hash to the Git commit so git checkout + dvc checkout restores both.
Reproducible Pipelines
DVC pipelines declare stages with inputs and outputs in dvc.yaml. DVC only reruns a stage when its dependencies change, giving cached, deterministic runs.
stages:
prepare:
cmd: python src/prepare.py
deps: [data/raw.csv, src/prepare.py]
outs: [data/clean.csv]
train:
cmd: python src/train.py
deps: [data/clean.csv, src/train.py]
params: [train.n_estimators, train.max_depth]
outs: [models/model.pkl]
metrics: [metrics.json]
Run the whole DAG with dvc repro and compare experiments with dvc exp show.
Model Packaging & Serving
Wrap the model in a REST API so applications can request predictions over HTTP. FastAPI is the standard choice: fast, async, with automatic validation and OpenAPI docs.
from fastapi import FastAPI
from pydantic import BaseModel
import joblib
app = FastAPI(title="Churn API")
model = joblib.load("models/model.pkl")
class Features(BaseModel):
tenure: float
monthly_charges: float
contract_type: int
@app.post("/predict")
def predict(f: Features):
x = [[f.tenure, f.monthly_charges, f.contract_type]]
proba = float(model.predict_proba(x)[0][1])
return {"churn_probability": proba,
"prediction": int(proba > 0.5)}
@app.get("/health")
def health():
return {"status": "ok"}
Run locally with uvicorn app:app --reload. Now package it so it runs identically anywhere.
Dockerizing the Service
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
docker build -t churn-api:1.0 .
docker run -p 8000:8000 churn-api:1.0
CI/CD for ML
Continuous integration for ML runs data validation, unit tests, and (optionally) a training + evaluation gate on every push. A model is only promoted if it beats the current production baseline.
# .github/workflows/ci.yml
name: ml-ci
on: [push]
jobs:
test-and-train:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -r requirements.txt
- run: pytest tests/
- run: dvc pull && dvc repro
- run: python src/evaluate.py --min-accuracy 0.85
Model Registry
A model registry is the source of truth for which model version is in which stage (Staging, Production, Archived). MLflow's registry manages this transition lifecycle.
from mlflow import MlflowClient
client = MlflowClient()
mv = mlflow.register_model("runs:/<run_id>/model", "churn-model")
client.set_registered_model_alias(
"churn-model", alias="production", version=mv.version)
# Serving code loads by alias, not a hardcoded version
model = mlflow.pyfunc.load_model("models:/churn-model@production")
Monitoring & Drift Detection
Models degrade silently as the world changes. Two failure modes to watch: data drift (input distribution shifts) and concept drift (the relationship between inputs and target changes).
from evidently import Report
from evidently.presets import DataDriftPreset
report = Report(metrics=[DataDriftPreset()])
result = report.run(reference_data=train_df, current_data=live_df)
result.save_html("drift_report.html")
Track prediction latency, error rates, and feature statistics with Prometheus + Grafana. Trigger retraining when drift crosses a threshold or accuracy drops below the SLA.
Feature Stores
A feature store (e.g. Feast) centralizes feature computation so training and serving use identical logic, eliminating training-serving skew. It offers an offline store for batch training and an online store for low-latency inference.
Orchestration
Pipelines need scheduling, retries, and dependency management. Airflow and Prefect are the common orchestrators.
from prefect import flow, task
@task(retries=2)
def extract(): ...
@task
def train(data): ...
@flow(name="daily-retrain")
def pipeline():
data = extract()
train(data)
if __name__ == "__main__":
pipeline.serve(cron="0 2 * * *") # 2am daily
Maturity Ladder
Level 0 is manual notebooks. Level 1 automates the training pipeline. Level 2 adds full CI/CD with automated retraining, testing, and deployment. Advance one rung at a time rather than jumping to full automation prematurely.
Practice Exercises
- Wrap a scikit-learn training script in an MLflow run that logs at least three parameters, one metric, and the model artifact. Open the MLflow UI and compare two runs.
- Initialize DVC in a repo, track a CSV dataset, push it to a local remote, then simulate a teammate with
dvc pull. - Build a FastAPI service with
/predictand/healthendpoints, then containerize it with a Dockerfile and run the image. - Write a
dvc.yamlwith prepare and train stages and confirmdvc reprocaches unchanged stages. - Register a model in the MLflow registry, promote a version to the
productionalias, and load it by alias. - Generate an Evidently data-drift report comparing a reference set to a shifted current set. Which features drift most?