Cutting an observability bill by ~89% at 6 TB/day
The problem
At one terabyte of telemetry a day, the platform was paying $80,000 a month to watch its own systems — $40K to Datadog and $40K to AWS CloudWatch, running side by side, neither aware of the other.
That duplication was not incompetence. The application team had instrumented with Datadog. The infrastructure team had CloudWatch on by default, and a compliance requirement had been satisfied that way years earlier. Both solved their own problem correctly. Nobody had ever put the two invoices on one page — one arrived through procurement, the other was buried inside a seven-figure AWS bill as a line called CloudWatch.
Meanwhile the data volume was heading toward six terabytes a day. On a per-GB meter, that is a linear problem with no ceiling.
The approach
I replaced both systems with a self-hosted Grafana LGTM stack (Loki, Mimir, Tempo, Grafana, with Alloy agents) on Kubernetes — four EKS clusters, blue-green, zero downtime through the cutover. Six levers did the cost work:
- Tiered storage — 90 days hot in S3, everything older to Glacier Deep Archive. 96% cut on cold-log storage. Trade-off: retrieval takes hours, the right trade for 90-day-old logs.
- Single-AZ pipeline — cross-zone transfer is a tax on every internal hop. ~$1,800/month eliminated.
- ARM Graviton + Karpenter right-sizing — 40% compute cut, no performance loss and no architecture change.
- NVMe-backed memcached — 80% cache hit rate, roughly 5× less query load on S3.
- Real multi-tenancy — one platform serving 15+ teams isolated by tenant ID, instead of fifteen separate stacks each carrying its own idle headroom.
- Retention as a decision, not a default — what gets kept, at what class, for how long, agreed per tenant rather than inherited.
The numbers
| Monthly cost | |
|---|---|
| Received bill at 1 TB/day Datadog $40K + CloudWatch $40K |
$80,000real invoice |
| SaaS-equivalent at 6 TB/day, 90-day retention Datadog ~$120K + CloudWatch ~$105K |
~$225,000projected — see note |
| Self-hosted LGTM at 6 TB/day, 90-day retention | ~$25,000real operated cost |
A per-GB SaaS meter scales linearly against you — every unit of product success bills you more to watch it — while a stack you own scales on compute and storage, both of which the six levers bend.
What broke
No migration this size is clean. The things that actually cost time:
- Single-AZ is a real trade, not a free win. It removes cross-zone cost and it removes zone-level redundancy with it. I bought the resilience back with blue-green secondary clusters and documented failover — engineering work that has to be budgeted, not a footnote.
- Karpenter provisioning failures under load. Pods stuck Pending when node-pool limits were hit or subnet IP capacity ran out. Aggressive right-sizing means you meet your capacity ceilings, and you find them the hard way first.
- Cache performance degraded quietly. When the memcached hit rate slipped below ~70%, queries slowed and S3 request volume climbed — a cost regression that shows up as a latency complaint. It needed replica scaling and NVMe sizing as a standing operational concern.
- Multi-tenancy is a discipline, not a config flag. A missing tenant header is an immediate hard failure. Every onboarding team hit it once.
- Upgrades had no automated rollback in the first phase. Helm changes were a human-error surface until GitOps was put in front of them. If I ran this again, that comes first, not later.
The result
A single multi-tenant observability platform, running 6 TB/day at 90-day retention for ~$25K/month, serving 15+ teams, with new teams onboarding at near-zero marginal cost. Zero downtime through the cutover. All infrastructure reproducible, so the platform outlives the person who built it.
The $80,000 a month was a real invoice. The ~$25,000 is a real operated cost. Between those two figures the volume grew six times.
You can stop after any step, and if your bill is already lean I will tell you that and charge you nothing.
Book the free review →