Skip to content
Reliability & Integration

intermediate · Interactive lab

API Rate Limiting

HTTP 429 is not a fault report: it is the API telling you that you are sending too much. Learn to read your allowance, respect the signal, and pace work through a window.

By the end: Read an allowance: what has been spent in this window, what the last answer said, and what may be sent right now.

Your challenge

Start here. This lab opens with the usual reflex, and it is wrong: every request is sent the moment it is ready, and a 429 is treated as a failure to retry at once. That spends the allowance on refused requests and turns a short pause into a long one. Read each moment, choose, and test.

See it happen

Watch the window fill

Caller: within its share
Slowed down so you can follow
READYREQUESTSWORK QUEUE0Integration LogicWHAT WE THINK WE HAVE SPENTWarehouse APIWHAT IT HAS REALLY COUNTED
Work arrives in a queue at Integration Logic, which sends it to the Warehouse API. The API allows a number of requests per window, answers 429 when the caller goes over, and sometimes says how long to wait before sending again.
  1. Work queue → Integration Logic
  2. Integration Logic → Warehouse API: within the allowance
STEP 1 / 6

An allowance, not an error budget

Integration Logic has work to push to the Warehouse API, which allows 10 requests every 10 seconds. Watch the caller's own count on the left and what the API is holding on the right.

Configure

Your sending decisions

For each request, choose what Integration Logic should do with it. Read its own count and the API's last answer first, and follow the API when the two disagree.

What should Integration Logic do with this request?
Send it now
· Spend one of the allowance on it immediately.
Pace the queue
· Send at the rate the window can carry, not all at once.
Wait for the window
· Hold until the caller's allowance refills.
Wait as the API asked
· Hold for exactly the Retry-After it sent.
Shed it
· Drop the work: it will be worthless by the time it could go.
  1. RQ-901 — Room left, one request

    A single stock update, and the window has barely been touched.

    Our count
    2 of 10 used · 7 s left in this window
    Last answer
    200 OK
    Allowed now
    8 requests before the limit
    The work
    1 request · a customer is waiting
  2. RQ-902 — One left, a queue behind it

    The nightly price refresh has 40 updates ready at once, and the allowance has almost run out.

    Our count
    9 of 10 used · 6 s left in this window
    Last answer
    200 OK
    Allowed now
    1 request before the limit
    The work
    40 requests · nightly batch, due within the hour
  3. RQ-903 — Allowance spent

    Everything this window allows has been sent, and the window has not rolled over yet.

    Our count
    10 of 10 used · 3.5 s left in this window
    Last answer
    200 OK
    Allowed now
    Nothing until the window refills in 3.5 s
    The work
    12 requests · nightly batch, due within the hour
  4. RQ-904 — The API said how long

    Our own count says there is room. The Warehouse API disagrees, and it is the one keeping score.

    Our count
    4 of 10 used · 4 s left in this window
    Last answer
    429 · Retry-After 20 s, 18 s of it still to run
    Allowed now
    Nothing for another 18 s, because the API said so
    The work
    12 requests · nightly batch, due within the hour
  5. RQ-905 — A 429 with no hint

    No Retry-After came with it, so there is nothing to obey — only the caller's own window to fall back on.

    Our count
    6 of 10 used · 5 s left in this window
    Last answer
    429 · no Retry-After given
    Allowed now
    4 requests before the limit
    The work
    12 requests · nightly batch, due within the hour
  6. RQ-906 — Inside the pause it asked for

    Sixteen seconds of the pause are still to run, and the caller's window looks as though it refills in two.

    Our count
    10 of 10 used · 2 s left in this window
    Last answer
    429 · Retry-After 20 s, 16 s of it still to run
    Allowed now
    Nothing for another 16 s, because the API said so
    The work
    12 requests · nightly batch, due within the hour
  7. RQ-907 — A long pause, perishable work

    This is a live price for a checkout page. It is replaced every 30 seconds, so a copy sent later is a copy nobody wants.

    Our count
    10 of 10 used · 5 s left in this window
    Last answer
    429 · Retry-After 45 s, 45 s of it still to run
    Allowed now
    Nothing for another 45 s, because the API said so
    The work
    1 request · worthless after 30 s

Two moments can show the same count and need different answers: it depends on what the API last said, and on what the work is.

Learn more

Why this pattern exists

The Warehouse API allows Integration Logic 10 requests every 10 seconds. Nothing is broken when the eleventh comes back 429: the API is telling the caller that the caller is sending too much. Integration Logic keeps its own count of what it has spent in the current window, and that count can be wrong — clocks drift, windows do not line up, and other callers may be spending the same quota. So when the API sends a Retry-After, that number beats the caller's own arithmetic. And on this API, as on many, a refused request still counts: pushing harder into a limit makes the limit tighter.

Integration Logic pushes work to the Warehouse API, which allows 10 requests every 10 seconds and sometimes says how long to wait. Each graded case is one moment: what Integration Logic had counted, what the API last answered, and what the request is for. What the API would really do with a request is hidden until one is sent.

  • Read an allowance: what has been spent in this window, what the last answer said, and what may be sent right now.
  • Tell the API's own signal apart from the caller's arithmetic, and follow the signal when they disagree.
  • Spread work across windows instead of bursting into them, and know when waiting is no longer worth it.

The rule this lesson applies: A rate limit is an allowance, not an error. Two habits waste it. The first is bursting: pushing a queue of work as fast as the code can loop, which empties the window in a second and turns the rest into 429s — and since refused requests are still counted here, the caller has paid for nothing. The second is retrying inside a pause the API asked for, which is how a short Retry-After becomes a long one. Pacing is the opposite habit: send at the rate the window can carry, so the quota is spent on work that is actually delivered. Waiting is not free either, though. Quota left unspent is throughput lost, and work that goes stale while the caller waits is worse than work never attempted — at some point the honest answer is to shed it. Rate limiting and circuit breaking are not the same instrument: a breaker stops calls to a dependency that is unwell; a rate limit is a healthy dependency telling a caller what its share is.