learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Idempotency & Retries

A payment whose answer is lost: why a blind retry charges twice, how an Idempotency-Key makes the retry safe, which failures are worth retrying, how retries at every layer multiply into 27 requests, and how jitter and retry budgets keep a recovering service alive.

An interactive System Design lesson: 19 steps, about 32 minutes, on a live simulation in your browser.

A checkout page, the client, charges a customer 42.00: POST /v1/charges to the payments service, 20 ms away. Payments charges the card (ch_1) and answers 201 Created. The answer is lost on the way back: a dropped connection, a load balancer restart, a mobile network changing cells.

The client waits its 1 s timeout, hears nothing and reports 504 Gateway Timeout. The customer sees an error. The card was charged.

What you will learn

  1. The ambiguous timeout

    • A charge whose answer is lost: A timeout does not mean the request failed; it means you do not know whether it did.
    • Retry it: Retrying a request that is not idempotent turns one lost answer into a duplicate action.
  2. Idempotency keys

    • Keys on the server only: Idempotency is a contract with two halves: the server remembers keys, and the client sends the same key on every attempt of one intent.
    • The client sends a key: With an idempotency key a retry asks "what happened to this payment?" instead of placing it again.
    • A retry while the first still runs: A key has three states: unknown (run it), in progress (409, try later), done (replay the stored answer).
    • One key, two different payments: A key names one exact request: the same key with a different body is a client bug, refused with 422 and never retried.
    • Drill: the header
    • Keys do not live for ever: A key protects only while it is remembered: its TTL must outlast every client's retries.
  3. What to retry, and when

    • 429 with Retry-After: Retry what might succeed next time (timeouts, 409, 429, 5xx), never what will fail the same way (400, 404, 422), and wait at least as long as Retry-After says.
    • Two clocks on every retry: Every retry policy needs two clocks: a timeout per attempt and a deadline for the whole call that the retries cannot outlive.
  4. Retries multiply

    • Retries at every layer: Retries at n layers multiply: 3 attempts at each of 3 layers is 3 × 3 × 3 = 27 requests for one call.
    • Retry at one layer: Retry at one layer, the one nearest the user, and let every layer below it fail fast.
    • Drill: the multiplier
  5. The thundering herd

    • 200 clients retry in step: Clients that fail together and back off by the same formula retry together: backoff alone moves the spike, it does not flatten it.
    • Full jitter: Jitter spreads synchronised retries into a trickle the recovering service can actually serve.
    • A retry budget fails fast: A retry budget limits retries to a share of traffic: when most calls fail, the caller stops retrying and fails fast.
    • Drill: the backoff
  6. Recap & playground

    • Cheat sheet
    • Playground