Skip to content
Reliability & Integration

beginner · Interactive lab

Retry Decisions

The same HTTP status can require different actions. Learn to decide from the failure context — not from a cheat sheet.

By the end: Read a failure together with its context: who is waiting, how long the problem lasts, and whether anything can be corrected automatically.

Your challenge

Start here. This lab opens with a status-code cheat sheet: every 503 is retried, every 401 refreshes the token, every 429 waits, and every 400 is parked. Some of those are wrong for their situation. Read each failure, choose, and test.

See it happen

Watch five failures, then decide

Warehouse healthy
Slowed down so you can follow
REQUESTANSWERPARKED FOR A PERSONOrder ProcessingWarehouse API● HEALTHYDLQIntegration LogicREADS CODE + CONTEXT
Order Processing sends an order to Integration Logic, which calls the Warehouse API. When the Warehouse API answers with an error, Integration Logic decides: retry, correct and retry, wait, park the order in the DLQ, or fail fast back to Order Processing.
  1. Order Processing → Integration Logic
  2. Integration Logic → Warehouse API: delivery attempt
  3. Integration Logic → DLQ: parked for a person
STEP 1 / 6

One question, many answers

Order Processing sends an order. Integration Logic calls the Warehouse API. When the Warehouse API answers with an error, the status code says what went wrong — not what to do. For that, Integration Logic reads the context.

Configure

Your decisions

For each failure, choose what Integration Logic should do. Read the whole situation, not just the code.

What should Integration Logic do?
Retry same request
· Send it again unchanged after a short backoff.
Retry after correction / auth refresh
· Correct the data or renew the credential, then send it.
Wait and respect Retry-After
· Wait as long as the API said, then send it.
Park / quarantine
· Keep it in the DLQ until a person fixes it.
Fail fast
· Tell the caller now that it did not go through.
  1. ORD-301 · HTTP 503

    The Warehouse API is overloaded for a few seconds and already recovering. This is the nightly stock sync: nobody is waiting.

  2. ORD-302 · HTTP 503

    Planned maintenance for the next two hours. A customer is waiting at checkout, which must answer within 30 seconds, and no order may be confirmed without the warehouse.

  3. ORD-303 · HTTP 401

    The access token expired a minute ago. The refresh token is still valid, so a new access token can be fetched automatically.

  4. ORD-304 · HTTP 401

    The warehouse team revoked this integration's credentials after a security review, and refreshing is refused too. Nightly batch: nobody is waiting.

  5. ORD-305 · HTTP 429

    Too many requests. The response says Retry-After: 20 seconds. It is a stock update that is due within the hour.

  6. ORD-306 · HTTP 429

    The daily request quota is used up: Retry-After is 12 hours. A customer is waiting at checkout, which must answer within 30 seconds.

  7. ORD-307 · HTTP 400

    The API rejected the country "India"; it accepts only ISO codes such as "IN". Integration Logic's country lookup maps one to the other.

  8. ORD-308 · HTTP 400

    The order names a product code the warehouse has never stocked, and no rule can say which product was meant. Nightly batch: nobody is waiting.

One decision per failure. Two failures can share a code and still need different decisions.

Learn more

Why this pattern exists

When the Warehouse API answers with an error, Integration Logic has to choose what happens to the order. Retrying is only one choice. Some failures clear on their own; some need something corrected first; some come with an instruction to wait; some need a person; and sometimes the most useful thing is to say no at once. The status code tells you what kind of problem the API saw. It does not tell you which of these to do.

Order Processing sends orders through Integration Logic to the Warehouse API. Each graded failure is a fixed, authored situation: an HTTP status code plus the context Integration Logic can see. No real API is called.

  • Read a failure together with its context: who is waiting, how long the problem lasts, and whether anything can be corrected automatically.
  • Choose between retrying the same request, correcting it first, waiting for Retry-After, parking it for a person, and failing fast.
  • See why the same status code can call for different decisions.

The rule this lesson applies: Retry the same request only when time will fix the problem and the request itself is fine. Correct the request or refresh the credential first when a known rule can fix it. When the API sends Retry-After, respect it, unless someone is waiting who cannot wait that long. Park the order in the DLQ when only a person can fix it and nobody is waiting. Fail fast when someone is waiting and no retry can succeed in time. This lesson grades every failure against its whole context, never the status code alone.