CASE STUDY

Azure FinOps
Platform

Two clouds. One question: who should be charged?

A production cost platform for a dual-cloud Azure environment: commercial and government tenants, dozens of subscriptions, and AI spend nobody could see. Built to answer what is being spent, by whom, and who should be charged for it, and live in front of leadership within days of the first working day. It has since grown into a fourteen-page platform whose newest wing turns visibility into action.

8 days
first working day to live in production
44
subscriptions across two cloud tenants
52→22%
unclassified spend, in a single day
13 mo
of billing history backfilled
01

The problem

Cloud spend had climbed roughly 75% in five months and nobody could say precisely why. The environment spanned two Azure tenants that behaved differently and were being discussed as one number. Tagging was jumbled and unclear: ten allocation concepts spelled twenty-five different ways, which meant finance had cost centers and project codes on paper and no defensible way to charge anything back. And a fast-growing slice of the bill, AI coding assistants, was invisible: buried inside the cloud invoice and double-counted in one briefing before this platform existed to say otherwise.

02

The system

The platform reads the cloud's own cost APIs directly: no export pipelines, no data factory, nothing to babysit. Every row carries both billed and effective cost side by side, FOCUS-style, so a forgotten filter can't silently double a total. Resource tags are ingested at the grain where they actually live and resolved retroactively, which is how a year of unclassified spend got reclassified in an afternoon without collecting a single new byte. AI assistant spend is ingested from the source billing API and reconciles to the vendor's own export to the cent. And the platform practices what it preaches: it runs on serverless infrastructure tuned so hard that its own bill is about $20 a month. Even the sync schedule is cost-engineered: the daily ingest jobs ride inside a window the platform already keeps awake, which makes the identical work roughly thirteen times cheaper than a midnight cron would be.

03

What it surfaced

THE HEADLINE

Essentially 0% Commitment Coverage

Billed versus effective cost differed by pocket change across the whole environment: nearly every dollar was running at on-demand rates. On a flat bill that's defensible; on one that grew 75% in five months it's the largest untouched lever, and now it's quantified, owned, and on the roadmap instead of invisible.

AI SPEND

A 5x Growth Curve Nobody Saw

AI assistant spend had quintupled in twelve months inside the cloud bill. The backfill made it visible, attributed it to cost centers and named users, and corrected a five-figure double-count before it reached the CFO twice.

OWNERSHIP

Spend with No Owner, Named

Ownerless license seats, untraced developer tooling, and an uncapped pay-as-you-go AI service that went from zero to real money in three weeks: each surfaced, quantified, and routed to an owner or explicitly accepted.

GOVERNANCE

A Tagging Standard That Stuck

Three required tags, one grain, ratified across finance, IT, and research. Charge-code coverage more than tripled in a day, retroactively across thirteen months, because the fix was reading existing tags correctly rather than demanding new ones.

04

The Optimization Engine

The platform's newest wing closes the loop from seeing spend to changing it, across five tabs: engineering actions, commitments, rightsizing, waste, and environment health. It reads the cloud's own advisor alongside utilization the platform measures itself, and it holds every recommendation to the same standard of evidence as the cost data.

LIFECYCLE

Realized Means the Bill Went Down

Every recommendation moves through an append-only lifecycle, and "realized" is never self-declared: a saving counts only after five settled billing days hold below 70% of the pre-decision baseline. No slide-deck savings, no double counting, and a permanent record of what was proposed, decided, and proven.

RIGHTSIZING

Sizing That Respects Reality

Utilization is banded by spikiness, a downsize must clear both CPU and memory at one in-family step down, savings are priced at a 70% recovery discount, and workload types get their own calibrated thresholds instead of one-size-fits-none.

WASTE

Six Classes of Waste, Named

Log ingestion, machines stopped but still billing, unattached disks, aging snapshots, empty service plans, and cold storage on hot tiers: each class quantified with evidence and routed to an owner, none hand-waved.

05

The quality practice

Cost data fails quietly: loaders that report success while inserting nothing, APIs that cap responses and return 200, gaps that hide behind a true "loaded through" date. This platform treats that as the primary threat. Invariant checks assert properties that must hold no matter which bug appears next; every gap detector must be proven able to fail before it's trusted; provisional data is labeled rather than hidden.

"A check that has never failed is not yet a check."

from the platform's engineering log

06

The stack

AzureAzure GovernmentCost Management APIFOCUS Resource GraphAzure SQL serverlessNodeReact Entra IDGitHub billing APIApp ServiceApplication Insights
07

Contact