← All posts · Incident · January 18, 2024 · 5 min read
Postmortem: the eu-central rebalance
On January 16 our query p99 in eu-central rose from 41ms to just over 900ms for 34 minutes. No data was lost and no alerts were missed. This is what happened and what we changed.

Timeline
- 22:03 UTC — a routine storage rebalance began moving shards between nodes.
- 22:09 — one node ended up owning two of the busiest shards at once. Its p99 climbed and the anomaly detector paged the primary on-call.
- 22:21 — we paused the rebalance and manually moved one hot shard to an idle node.
- 22:37 — latency returned to baseline. The rebalance resumed with the new limits below.
Root cause
The rebalancer optimised for even shard count per node, not for the load behind each shard. Two low-numbered shards happened to carry a disproportionate share of traffic, and nothing stopped them from landing together.
What we changed
The rebalancer now weights shards by recent read load and refuses to co-locate two shards above a load threshold. We also added an SLO burn-rate alert that fires on the rate of latency increase, so the next one pages us a few minutes earlier.
No customer data was affected at any point, and retention windows were untouched.