What FinOps Actually Is
FinOps is the operating model for spending cloud money well — a collaboration between engineering, finance, and product to make cost a first-class metric alongside performance and reliability. Its lifecycle runs in three repeating phases: Inform (visibility, allocation, benchmarking), Optimize (rightsizing, commitments, waste removal), and Operate (governance, forecasting, and continuous improvement).
The core insight
Cloud turns capex into opex and pushes purchasing decisions down to the engineers who deploy resources. The team that spins up an oversized instance is the team best positioned to right-size it. FinOps makes that spend visible to them in near-real time so they can act.
Pricing Models
Every optimization decision starts with understanding how you are billed. The same VM can cost wildly different amounts depending on the purchasing model.
| Model | Discount | Commitment | Best for |
|---|---|---|---|
| On-Demand | 0% (baseline) | None | Spiky, unpredictable workloads |
| Spot / Preemptible | Up to ~90% | None (can be reclaimed) | Fault-tolerant batch, CI, big data |
| Savings Plans / CUDs | Up to ~72% | 1 or 3 yr spend/usage | Steady baseline compute |
| Reserved Instances | Up to ~72% | 1 or 3 yr, specific family | Stable, known instance types |
| Concept | AWS | Azure | GCP |
|---|---|---|---|
| Cheap interruptible | Spot Instances | Spot VMs | Spot / Preemptible VMs |
| Commitment discount | Savings Plans / RIs | Reservations / Savings Plans | Committed Use Discounts |
| Cost analysis | Cost Explorer | Cost Management | Cloud Billing / BigQuery export |
| Budgets & alerts | AWS Budgets | Cost alerts | Budgets & alerts |
Visibility: Tagging & Allocation
You cannot optimize what you cannot attribute. A consistent tagging strategy (or labels/resource groups) lets you split the bill by team, environment, and product. Enforce required tags at creation time and treat untagged spend as a defect.
# Tag resources for cost allocation
aws ec2 create-tags --resources i-0abc123 \
--tags Key=team,Value=payments Key=env,Value=prod Key=cost-center,Value=CC42
# Activate a tag key as a cost allocation dimension
aws ce list-cost-allocation-tags --status Active
# GCP labels
gcloud compute instances update my-vm --zone=us-central1-a \
--update-labels=team=payments,env=prod
Querying spend
# AWS: month-to-date cost grouped by service
aws ce get-cost-and-usage \
--time-period Start=2026-06-01,End=2026-07-01 \
--granularity MONTHLY --metrics UnblendedCost \
--group-by Type=DIMENSION,Key=SERVICE
# Azure: cost by resource group
az costmanagement query --type ActualCost \
--scope "/subscriptions/SUB" --timeframe MonthToDate \
--dataset-grouping type=Dimension name=ResourceGroupName
Optimize: Rightsizing & Waste
Most cloud bills contain 20-40% pure waste. The biggest wins are the least glamorous: turn off what nobody uses. Hunt for idle instances, unattached disks, orphaned IPs and load balancers, over-provisioned instances, and non-prod environments running 24/7.
- Rightsize: match instance size to observed CPU/memory (target sustained utilisation, not peak of peaks).
- Schedule: stop dev/test environments nights and weekends — a 12x5 schedule cuts their cost ~65%.
- Clean up: delete unattached volumes, old snapshots, and idle elastic IPs.
- Tier storage: move cold objects to cheaper classes with lifecycle rules.
- Modernize: serverless and Arm/Graviton instances often cut cost per request substantially.
# Find unattached EBS volumes (pure waste)
aws ec2 describe-volumes --filters Name=status,Values=available \
--query "Volumes[].{ID:VolumeId,GB:Size,AZ:AvailabilityZone}" --output table
# Pull AWS-generated rightsizing recommendations
aws ce get-rightsizing-recommendation --service AmazonEC2
# Schedule a dev VM to stop nightly (GCP resource policy)
gcloud compute resource-policies create instance-schedule dev-off \
--region=us-central1 --vm-stop-schedule="0 20 * * 1-5" --timezone=UTC
Commit: Savings Plans & Reservations
Once you know your steady-state baseline, cover it with a commitment. Savings Plans (AWS) commit to a dollar-per-hour of compute spend and flex across instance families and regions; Committed Use Discounts (GCP) and Reservations (Azure) are the equivalents. The rule of thumb: cover roughly 60-80% of your predictable baseline with commitments and let the variable top layer run on-demand or spot.
Don't over-commit
A 3-year commitment you cannot fully use is worse than paying on-demand — you pay for idle capacity for years. Start with 1-year, no-upfront Savings Plans covering only the workload you are confident stays flat, then layer more as confidence grows. Track your commitment utilisation and coverage as first-class KPIs.
Operate: Budgets, Anomalies & Forecasting
Governance keeps optimizations from decaying. Set budgets with alerts at 50/80/100% of forecast, enable anomaly detection so a runaway resource pages you the same day rather than at month-end, and review a unit-economics metric (cost per customer, per request) that ties spend to business value.
# Create a monthly budget with an 80% alert (AWS)
aws budgets create-budget --account-id 111122223333 \
--budget '{"BudgetName":"prod-monthly","BudgetLimit":{"Amount":"20000","Unit":"USD"},"TimeUnit":"MONTHLY","BudgetType":"COST"}' \
--notifications-with-subscribers '[{"Notification":{"NotificationType":"ACTUAL","ComparisonOperator":"GREATER_THAN","Threshold":80},"Subscribers":[{"SubscriptionType":"EMAIL","Address":"finops@acme.com"}]}]'
# Turn on cost anomaly detection
aws ce create-anomaly-monitor \
--anomaly-monitor '{"MonitorName":"acme-monitor","MonitorType":"DIMENSIONAL","MonitorDimension":"SERVICE"}'
Practice Exercises
- Define a tagging policy with three required tags (team, env, cost-center), apply it to a resource, and produce a cost report grouped by one of those tags.
- Use the cost analysis CLI to list your top five services by month-to-date spend, and explain which is a candidate for optimization and why.
- Find all unattached volumes and idle IPs in an account, estimate their monthly cost, and write the cleanup commands (do not run destructively without review).
- Given a workload with a flat 40-vCPU baseline and spiky peaks to 100 vCPU, design a mix of Savings Plans, on-demand, and spot. Justify the coverage percentage you chose.
- Create a monthly budget with alerts at 80% and 100% of forecast delivered to email, then enable anomaly detection on the service dimension.
- Pick a non-prod environment and design a stop/start schedule. Calculate the percentage saving versus running 24/7 and describe how you would automate it.