Idempotency & Retries
A payment whose answer is lost: why a blind retry charges twice, how an Idempotency-Key makes the retry safe, which failures are worth retrying, how retries at every layer multiply into 27 requests, and how jitter and retry budgets keep a recovering service alive.
An interactive System Design lesson: 19 steps, about 32 minutes, on a live simulation in your browser.
A checkout page, the client, charges a customer 42.00: POST /v1/charges to the payments service, 20 ms away. Payments charges the card (ch_1) and answers 201 Created. The answer is lost on the way back: a dropped connection, a load balancer restart, a mobile network changing cells.
The client waits its 1 s timeout, hears nothing and reports 504 Gateway Timeout. The customer sees an error. The card was charged.
What you will learn
The ambiguous timeout
- A charge whose answer is lost: A timeout does not mean the request failed; it means you do not know whether it did.
- Retry it: Retrying a request that is not idempotent turns one lost answer into a duplicate action.
Idempotency keys
- Keys on the server only: Idempotency is a contract with two halves: the server remembers keys, and the client sends the same key on every attempt of one intent.
- The client sends a key: With an idempotency key a retry asks "what happened to this payment?" instead of placing it again.
- A retry while the first still runs: A key has three states: unknown (run it), in progress (409, try later), done (replay the stored answer).
- One key, two different payments: A key names one exact request: the same key with a different body is a client bug, refused with 422 and never retried.
- Drill: the header
- Keys do not live for ever: A key protects only while it is remembered: its TTL must outlast every client's retries.
What to retry, and when
- 429 with Retry-After: Retry what might succeed next time (timeouts, 409, 429, 5xx), never what will fail the same way (400, 404, 422), and wait at least as long as Retry-After says.
- Two clocks on every retry: Every retry policy needs two clocks: a timeout per attempt and a deadline for the whole call that the retries cannot outlive.
Retries multiply
- Retries at every layer: Retries at n layers multiply: 3 attempts at each of 3 layers is 3 × 3 × 3 = 27 requests for one call.
- Retry at one layer: Retry at one layer, the one nearest the user, and let every layer below it fail fast.
- Drill: the multiplier
The thundering herd
- 200 clients retry in step: Clients that fail together and back off by the same formula retry together: backoff alone moves the spike, it does not flatten it.
- Full jitter: Jitter spreads synchronised retries into a trickle the recovering service can actually serve.
- A retry budget fails fast: A retry budget limits retries to a share of traffic: when most calls fail, the caller stops retrying and fails fast.
- Drill: the backoff
Recap & playground
- Cheat sheet
- Playground