Cloud spend cut 44% while traffic grew
Per-customer cost attribution, a storage tier redesign, and an inference caching layer that turned an unpredictable AI bill into a forecastable line item.

Where they were
A Series B analytics platform was burning $310k/month in cloud spend with no idea which customers or features drove it.
A newly launched AI feature had added $70k/month of model cost in one quarter, growing faster than the revenue attached to it.
The CFO wanted a number for the board; nobody could produce one.

What we did
01Attribution before optimization
Tagging and instrumentation to get cost per customer, per feature, and per request. Two enterprise accounts turned out to be gross-margin negative.
02Right-sized the obvious things
Idle non-production environments, over-provisioned instances, and eleven months of logs in hot storage. Unglamorous, and it paid for the engagement in week two.
03Redesigned the storage tiers
Query patterns showed 91% of reads hit the last 30 days. Older data moved to columnar cold storage with transparent access, cutting storage cost by two-thirds.
04Made AI cost a design parameter
Semantic caching, prompt compaction, and tiered model routing — cheap model first, escalate only on low confidence. Same output quality on the eval set at 38% of the token cost.
“The attribution work was the real deliverable. The savings were great, but knowing which customers cost us money changed how we price.”
What it was built with
Let’s talk about yours.
Mention this case study in your note and we’ll come to the call with the specifics of how it went — including what we’d do differently.