Cutting an observability bill by ~89% at 6 TB/day

Role: engineering manager, a high-scale real-money gaming platform (anonymised)
What I did: architect and operator — I built the replacement and ran it
Scope: metrics, logs and traces for 15+ engineering teams

The problem

At one terabyte of telemetry a day, the platform was paying $80,000 a month to watch its own systems — $40K to Datadog and $40K to AWS CloudWatch, running side by side, neither aware of the other.

That duplication was not incompetence. The application team had instrumented with Datadog. The infrastructure team had CloudWatch on by default, and a compliance requirement had been satisfied that way years earlier. Both solved their own problem correctly. Nobody had ever put the two invoices on one page — one arrived through procurement, the other was buried inside a seven-figure AWS bill as a line called CloudWatch.

Meanwhile the data volume was heading toward six terabytes a day. On a per-GB meter, that is a linear problem with no ceiling.

The approach

I replaced both systems with a self-hosted Grafana LGTM stack (Loki, Mimir, Tempo, Grafana, with Alloy agents) on Kubernetes — four EKS clusters, blue-green, zero downtime through the cutover. Six levers did the cost work:

  1. Tiered storage — 90 days hot in S3, everything older to Glacier Deep Archive. 96% cut on cold-log storage. Trade-off: retrieval takes hours, the right trade for 90-day-old logs.
  2. Single-AZ pipeline — cross-zone transfer is a tax on every internal hop. ~$1,800/month eliminated.
  3. ARM Graviton + Karpenter right-sizing — 40% compute cut, no performance loss and no architecture change.
  4. NVMe-backed memcached — 80% cache hit rate, roughly 5× less query load on S3.
  5. Real multi-tenancy — one platform serving 15+ teams isolated by tenant ID, instead of fifteen separate stacks each carrying its own idle headroom.
  6. Retention as a decision, not a default — what gets kept, at what class, for how long, agreed per tenant rather than inherited.

The numbers

Monthly cost 
Received bill at 1 TB/day
Datadog $40K + CloudWatch $40K
$80,000real invoice
SaaS-equivalent at 6 TB/day, 90-day retention
Datadog ~$120K + CloudWatch ~$105K
~$225,000projected — see note
Self-hosted LGTM at 6 TB/day, 90-day retention ~$25,000real operated cost
The volume grew 6×. The bill did not follow it.
~89% reduction — about 9× cheaper — against the like-for-like SaaS setup.

A per-GB SaaS meter scales linearly against you — every unit of product success bills you more to watch it — while a stack you own scales on compute and storage, both of which the six levers bend.

How to read these figures. The $80,000 at 1 TB/day is an invoice that was actually received. The ~$25,000 is what the self-hosted stack actually cost to operate. The ~$225,000 is a projection — what that telemetry volume at that retention would have cost on those two vendors at 2024 list pricing. It is a like-for-like replacement comparison, not a bill anyone was sent, and I label it every time. A teardown that does not tell you which of its figures are invoices and which are models should be thrown away.

What broke

No migration this size is clean. The things that actually cost time:

The result

A single multi-tenant observability platform, running 6 TB/day at 90-day retention for ~$25K/month, serving 15+ teams, with new teams onboarding at near-zero marginal cost. Zero downtime through the cutover. All infrastructure reproducible, so the platform outlives the person who built it.

The $80,000 a month was a real invoice. The ~$25,000 is a real operated cost. Between those two figures the volume grew six times.

Free 15-min bill review → Cloud Bill Teardown → Cloud Escape Rescue → managed platform

You can stop after any step, and if your bill is already lean I will tell you that and charge you nothing.

Book the free review →
Rajesh Medampudi · [email protected] · rajesh.medampudi.com · cal.com/rajesh-medampudi/bill-review