Insights · Azure & AWS FinOps, from the people who build the checks
Deep, practical guides on finding and eliminating waste on Azure and AWS — written by the team behind the CloudFinOpsKit tools' cost checks. Real meter behaviour, real CLI commands, real traps we've hit so you don't have to.
A sudden spike, investigated in order: confirm it's real (unfinalized days, commitment purchases, expired credits), isolate the service and the day it started, drill to the resource, match it to the usual suspects — data-transfer blow-outs, scale-outs, snapshot sprawl, leaked keys — and stop the bleed. With the exact Cost Explorer and Cost Management steps.
A prioritized, do-it-in-order playbook: kill zero-risk waste first, turn off non-prod out of hours, rightsize, apply Hybrid Benefit, layer reservations & savings plans, tame Log Analytics, and make it a monthly habit — each step with the exact Azure action.
An honest comparison: the free AWS-native tools (Cost Explorer, Compute Optimizer, Cost Optimization Hub), commitment automation, CloudFinOpsKit, and enterprise platforms — what each does, what it costs, and who it's for.
Close the gap between what your pods request and what they use: right-size requests, bin-pack with autoscaling/Karpenter, run Spot, scale non-prod to zero, and cover the baseline with commitments.
Set up free, ML-based anomaly monitors, pick a threshold you can explain, read the root-cause hints, and pair it with budgets + a monthly review to catch slow creep.
Azure's anomaly alerts and AWS's Cost Anomaly Detection compared — scopes, tunability, delivery, blind spots — and one auditable deterministic methodology (same-weekday spikes, step changes, Theil–Sen creep, new/vanished spend) that gives both clouds a single definition of "anomaly".
Why token spend breaks classic anomaly detection, and the four detectors that work: tokens/day spikes, output:input ratio bloat (output costs 3–5× input), blended $/1K-token rate shifts that expose a model-mix change, and idle provisioned throughput.
The working checklist we run on AWS — kill waste (EBS, EIPs, idle NAT/LBs), right-size with Compute Optimizer, commit smart with Savings Plans, tier S3, govern with Tag Policies, and optimize Bedrock spend. Mapped to the Well-Architected Cost Optimization Pillar.
You can't manage what you don't meter. The economics half of tokenomics — meter token usage, price the true blended cost per query, attribute AI spend to an owner, and govern it as agents multiply the bill.
Tokens are the atomic unit of AI cost. Measure tokenomics, then pull four levers — right-size & route the model, trim context, cap output, cache repeats — to cut LLM spend 60-80% without losing quality.
AI cost is the #1 FinOps focus of 2026. Why it breaks the old playbook, the levers that control it (prompt caching, right model, PTU, output caps, zombies), and how to attribute token spend.
Showback vs chargeback, the three ways to allocate Azure spend, handling shared costs nobody owns, and producing a per-team statement finance will actually act on.
Built-in Cost Management anomaly alerts, the daily and month-over-month checks that matter, a threshold you can explain, and the usual suspects behind a spike.
Per-model token visibility from CloudWatch, then five levers — prompt caching, output caps, model right-sizing, context trimming, and the right Provisioned Throughput call — to cut Bedrock spend without losing quality.
The five tags that drive allocation, enforcing them with Azure Policy (require & inherit), backfilling an untagged estate, and tracking coverage as a KPI.
The preventive guardrails, detective review cadence and accountability loop that turn a one-off cleanup into a durable FinOps practice — plus a 90-day rollout.
The frontier of FinOps maturity — tie spend to a business outcome. Pick the right unit, compute cost per customer or transaction, and turn the metric into decisions.
Why every major cloud now emits FOCUS natively, the core columns, how to turn it on in Azure Cost Management, and why it makes your cost data AI-ready.
Find available EBS volumes with the Console, CLI and Cost Explorer — and dodge the traps: snapshots backing AMIs, volumes on stopped instances, idle Elastic IPs, and DR snapshots. With a safe, snapshot-first workflow.
Run-rate vs trend vs driver-based methods, the built-in Cost Management forecast, tracking budget variance to ±10%, and the commitment and anomaly factors that throw forecasts off.
Discount depth vs flexibility, use-it-or-lose-it utilization, exchange rules, the Dev/Test subscription myth, and the layering strategy mature teams use.
Portal, Resource Graph and CLI methods — plus the false positives that burn people: Site Recovery replica disks, backup snapshots, and disks on stopped VMs.
Compute SP vs EC2 Instance SP vs RIs — coverage, discount and the flexibility trade-off, how Graviton changes the math, and a coverage strategy that bounds stranded-commitment risk.
Who qualifies, who doesn't (Dev/Test subscriptions, Windows client VMs, SQL Express/serverless), the 8-core minimum, and the exact CLI to find every VM not using it.
Azure Cost Management, Advisor, Microsoft's FinOps toolkit, CloudFinOpsKit, and the enterprise platforms — what each actually does, what it costs, and when you've outgrown it.
The complete working checklist — compute, storage, network, databases, commitments, licensing, monitoring and AI workloads — grouped the way a real cost review runs.