Exponential backoff
Exponential backoff is a retry policy in which the wait between attempts grows by a constant factor each time, usually doubling.
Exponential backoff is a retry policy in which the wait between attempts grows by a constant factor each time, usually doubling. It gives a congested or rate-limited resource an increasing amount of time to recover instead of adding load to it.
#The three parts
All three are required. Growth alone is not a backoff policy.
delay = base * factor ** attempt # growth
delay = delay * (0.5 + random()) # jitter
delay = min(delay, ceiling) # bound
if attempt >= max_attempts: give_up() # termination
Jitter matters more than people expect. Parallel workers that fail together will retry together, reproducing exactly the burst that caused the failure, and a synchronised retry storm looks identical to an attack from the receiving end.
#Where the idea comes from
TCP has done this since long before HTTP clients did. RFC 6298 specifies that the retransmission timer is doubled on each timeout, that it should be rounded up to at least one second, and that any maximum imposed must be at least 60 seconds. RFC 8961 generalises the requirements for time-based loss detection. Borrowing a policy that the transport layer already applies underneath you is a reasonable default.
#When backoff is the wrong tool
| Response | Correct reaction |
|---|---|
| 429, or 503 with Retry-After | Wait as instructed, then back off |
| Connection reset or timeout | Back off; the cause may be transient |
| 403 | Change something. Repeating it unchanged will not succeed |
| A CAPTCHA interstitial | Not a rate problem. Retrying harder makes it worse |
An explicit Retry-After outranks your own calculation. Waiting less than the server asked is the most reliable way to convert a temporary limit into a durable block.
#Commonly confused with
Backoff is not a substitute for a correct request rate. If a steady-state rate is above what the target tolerates, backoff simply produces a slow oscillation around the limit and a poor success rate. Fix the rate, and see rate limiting and concurrency for the surrounding controls.