← All posts · Engineering · March 12, 2024 · 8 min read

Tail-based sampling in production

Head-based sampling decides whether to keep a trace the moment it starts — long before you know whether it turned slow or errored. Tail-based sampling waits until the trace is complete. That one change is what lets us store a small fraction of all traffic and still keep every interesting request.

Trace waterfall

Why head-based sampling loses the good ones

If you flip a coin at the first span, the slow checkout that happens once in ten thousand requests is almost always the one you dropped. You end up with a warehouse full of healthy 200s and none of the traces an on-call engineer actually needs at 3am.

Buffering a trace until it is done

Our collectors hold the spans of an in-flight trace in memory, keyed by trace ID, until the root span finishes or a timeout fires. Only then do the sampling rules run, with the whole trace in view: total duration, every status code, and which services it touched.

sampling:
  mode: tail
  decision_wait: 8s
  keep:
    - error == true
    - duration_ms > slo.p99
  probabilistic: 0.05

What it cost us

Buffering costs memory, not disk, and the memory is bounded by how many traces are open at once rather than by total volume. In exchange, trace storage dropped by roughly 90% while our "found the trace I needed" rate went up, because the 5% we keep is now the useful 5%.

Next: Taming metric cardinality →