intermediate · Interactive lab
API Rate Limiting
HTTP 429 is not a fault report: it is the API telling you that you are sending too much. Learn to read your allowance, respect the signal, and pace work through a window.
By the end: Read an allowance: what has been spent in this window, what the last answer said, and what may be sent right now.
Your challenge
Start here. This lab opens with the usual reflex, and it is wrong: every request is sent the moment it is ready, and a 429 is treated as a failure to retry at once. That spends the allowance on refused requests and turns a short pause into a long one. Read each moment, choose, and test.
See it happen
Watch the window fill

- Work queue → Integration Logic
- Integration Logic → Warehouse API: within the allowance
An allowance, not an error budget
Integration Logic has work to push to the Warehouse API, which allows 10 requests every 10 seconds. Watch the caller's own count on the left and what the API is holding on the right.
Learn more
Why this pattern exists
The Warehouse API allows Integration Logic 10 requests every 10 seconds. Nothing is broken when the eleventh comes back 429: the API is telling the caller that the caller is sending too much. Integration Logic keeps its own count of what it has spent in the current window, and that count can be wrong — clocks drift, windows do not line up, and other callers may be spending the same quota. So when the API sends a Retry-After, that number beats the caller's own arithmetic. And on this API, as on many, a refused request still counts: pushing harder into a limit makes the limit tighter.
Integration Logic pushes work to the Warehouse API, which allows 10 requests every 10 seconds and sometimes says how long to wait. Each graded case is one moment: what Integration Logic had counted, what the API last answered, and what the request is for. What the API would really do with a request is hidden until one is sent.
- Read an allowance: what has been spent in this window, what the last answer said, and what may be sent right now.
- Tell the API's own signal apart from the caller's arithmetic, and follow the signal when they disagree.
- Spread work across windows instead of bursting into them, and know when waiting is no longer worth it.
The rule this lesson applies: A rate limit is an allowance, not an error. Two habits waste it. The first is bursting: pushing a queue of work as fast as the code can loop, which empties the window in a second and turns the rest into 429s — and since refused requests are still counted here, the caller has paid for nothing. The second is retrying inside a pause the API asked for, which is how a short Retry-After becomes a long one. Pacing is the opposite habit: send at the rate the window can carry, so the quota is spent on work that is actually delivered. Waiting is not free either, though. Quota left unspent is throughput lost, and work that goes stale while the caller waits is worse than work never attempted — at some point the honest answer is to shed it. Rate limiting and circuit breaking are not the same instrument: a breaker stops calls to a dependency that is unwell; a rate limit is a healthy dependency telling a caller what its share is.

