Blog

Notes on Datadog cost overruns

Everything below comes from actually tracing real Datadog bills back to their cause — the patterns that kept showing up, and how to check for them in your own org.

Context

Why so many Datadog bills are exploding in 2026

28,000 mid-market Datadog customers have no dedicated FinOps team. Here's why that's exactly the segment where cost surprises happen.

Market context · 6 min read
Checklist

APM sample rate vs. Adaptive Sampling Target: the quiet cost killer

A tracer config that disagrees with your Datadog sampling target can push ingestion to 334% of budget without a single alert firing.

Detection checklist · 7 min read
Checklist

Retention filters stuck at 100%: how to check in five minutes

One API call tells you whether a retention filter is indexing everything, unfiltered, since the day it was created.

Detection checklist · 5 min read
Checklist

Synthetics, custom metrics, and the committed capacity nobody uses

Three unglamorous line items — polling synthetics, untagged custom metrics, unused committed capacity — and how to spot each one.

Detection checklist · 8 min read
Case study

A 43% cost spike, traced back to a single commit

Day-by-day usage history, walked backwards until the exact deploy date fell out of the data.

Case study · 6 min read
Walkthrough

Anatomy of a Datadog cost report

What a real diagnostic report contains, section by section — the summary, the cost-change table, the auto-drafted support tickets, and why the raw-findings appendix exists at all.

Walkthrough · 9 min read
Case study

The bill went down 44% the same month usage spiked 234%

Why a monthly rollup and a day-by-day usage chart can tell two different stories about the same product family — and why that's exactly the point of looking day by day.

Case study · 7 min read
Checklist

A retention filter at 100% that matches nothing

The inverse of the usual retention-filter problem: a filter costing zero in ingested volume, still sitting there as unmonitored config debt.

Detection checklist · 6 min read
Case study

When the tool can't tell you the answer

An Adaptive Sampling Target it couldn't match, a custom metric named "couter" — and why drafting the support-ticket question beats guessing.

Case study · 7 min read
Guide

Committed-use discounts: what to actually do about 100% on-demand spend

Eight product families, zero committed coverage, real dollar savings per family — and why the fix is a conversation with your account team, not a settings toggle.

Guide · 8 min read
Guide

12 months or longer: when a fixed plan actually pays off

A decision framework for committed-use plans — usage stability, marginal discount vs. lock-in length, and why not to commit everything at once.

Guide · 7 min read
Guide

CoreDNS and cert-manager are quietly flooding your custom metrics with untagged noise

39 untagged CoreDNS metrics, 7 from cert-manager — a pattern any Kubernetes cluster on Datadog can hit, and where the real fix lives.

Guide · 8 min read
Methodology

Ranked by dollar impact, not percent

Why a $977 on-demand exposure should outrank a dramatic-looking percentage on a $20 line item — and how the report's summary now sorts for that.

Methodology · 6 min read