postmortem · Sep 29, 2026
Postmortem: The 41-Minute Retry Storm
A 90-second provider blip turned into 41 minutes of our own errors. Timeline, the commands we ran, the three-line diff, and why a retry loop without jitter is a synchronized weapon.
Brindlecast fans notifications out to a mail provider. On the night of the incident the provider had a 90-second blip. Our error window lasted 41 minutes. The gap between those two numbers is this post.
Summary
For about 90 seconds, the provider answered our requests with timeouts. That was their problem and it was brief. Our problem began when it ended: all 200 delivery workers were retrying on the same 500 ms rhythm, so they hit the recovering provider in synchronized waves, and the provider's rate limiter answered the waves with rejections. We retried those too. We kept the incident alive ourselves for another 39 minutes.
Impact: delayed notifications for roughly 40 minutes, no lost messages, no data corruption. The numbers in this post describe an invented service and are illustrative.
Timeline
All times UTC.
Time | Event |
|---|---|
02:11 | Provider starts timing out. First 3 errors in the log. |
02:13 | Provider recovers. Our errors climb to nearly 10,000 a minute. |
02:19 | Error-rate alert fires; on-call acknowledges. |
02:27 | Workers restarted. Errors are back within 90 seconds. |
02:38 | Worker count cut from 200 to 20 by config. Errors fall. |
02:52 | Error rate normal. Incident closed. |
The restart at 02:27 is the telling line. If the provider had still been down, restarting wouldn't have helped. It was up. Our own behaviour was the outage.
What we ran
First question: when did the errors start, and when did they stop being the provider's fault? The service log has one line per event, like this:
2026-09-27 02:14:08 ERROR deliver failed status=429 attempt=3So one awk command can count error lines per minute:
awk '$3=="ERROR"{c[substr($2,1,5)]++} END{for(k in c)print k, c[k]}' relay.log | sort02:11 3
02:12 212
02:13 9870
02:14 23410
02:15 23060That's an excerpt, five minutes of a longer log. The shape is what mattered: the 212 in the second row are the provider timing out, and the tens of thousands after it are our workers being turned away. Errors kept rising after the provider's status page said it had recovered. A cause outside our service doesn't behave like that.
Root cause and the diff
The retry helper waited a fixed 500 ms between attempts. A fixed delay has no memory of who else is waiting, so every worker that failed at the same moment retried at the same moment, and at the moment after that. The helper after the fix doubles its ceiling on each attempt, caps it at 30 seconds, and picks a random wait below the ceiling:
@@ -1,3 +1,4 @@
-export function delays(attempts = 8) {
- return Array.from({ length: attempts - 1 }, () => 500);
+export function delays(attempts = 8, rand = Math.random) {
+ return Array.from({ length: attempts - 1 }, (_, i) =>
+ Math.round(rand() * Math.min(30000, 500 * 2 ** i)));
}Taking the random source as a parameter lets us test it. With rand pinned to 1 the ceilings are 500, 1000, 2000, 4000, 8000, 16000 and 30000 milliseconds. Pinned to 0, every wait is zero. The old helper returned 500 seven times. The randomness spreads 200 workers across the whole window, so the provider sees a trickle where it saw a wave.
What we're changing
Retry delays use exponential backoff with full jitter everywhere, not only here. Done this week.
A retry budget: each worker may spend at most 10 percent of its requests on retries. When the budget is gone, the message waits in the queue instead.
The alert now reports retries as a share of traffic, so the next storm shows up as a ratio before it shows up as errors.
A game day: block the provider for two minutes in staging and watch the recovery curve, not only the outage.
We were lucky that the failure was an outage and not a deploy. Either way, the fix is the same: when something comes back, don't all run at it at once.
No comments yet
Comments are open. Have a thought or a question? Share it below.