Retries, idempotency, and the double charge
The bug where a customer is billed twice, and the one idea that prevents it.
Networks fail in a particularly nasty way: the request arrives and succeeds, and then the response is lost on the way back. The caller cannot tell this apart from the request never arriving. Both look identical: silence.
The fix is idempotency: designing an operation so that performing it repeatedly has the same effect as performing it once. The usual mechanism is a key. The caller invents a unique id for the attempt and sends it with every retry. The server records which keys it has already processed and, on seeing a repeat, returns the original result instead of doing the work again.
- Attempt 1
Send with key abc-123. Charge happens. Response lost. - Attempt 2
Send again with the same key abc-123. - Server checks
Key abc-123 already processed. - Result
Original outcome returned. One charge. Caller is satisfied.
The other half of retrying well is backing off. Retrying immediately, repeatedly, against a struggling service is how a small problem becomes an outage. Everyone retries at once, the service gets more load precisely when it has least capacity, and it never recovers. Wait longer between each attempt, add a little randomness, and give up eventually.
If this runs twice, what does the user see?
Two emails is embarrassing. Two charges is a refund, an apology, and a support ticket. The cost of getting this wrong is not evenly distributed.
What to remember
- A lost response is indistinguishable from a failed request.
- Idempotency keys let a retry return the original result instead of repeating work.
- Retry with increasing delays and randomness, and stop eventually.
Terms in this lesson
Field notes
Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.
The digest that stopped for five weeks
A weekly email job silently stopped running after an infrastructure change. Nobody noticed, because a job that does not run produces no error. It was discovered when a customer asked whether they had been unsubscribed.
Why you alert on missing success
Everyone retried at once
A service had a brief wobble. Every client retried immediately, then again, then again. The retries were far more traffic than the original load, and the service never got a quiet moment to recover. The outage lasted forty minutes longer than the fault did.
The thundering herd
resolved in 900ms · region iad1
Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.