intermediate · Interactive lab
Agent Retries
A person retries once and thinks about it. A loop retries as fast as the network allows, several times, without pausing — so the retry policy is no longer advice, it is behaviour.
By the end: Read a repeat as a property of the tool: an idempotency key, a naturally idempotent call, or neither.
Your challenge
Start here. This lab opens with the loop retrying at every attempt, which is the wrong starting point and what an agent does by default: the retry is the cheapest thing it can do. One of these repeats may charge a customer twice, one goes into a pause the API asked for, one is past its budget, and one should be handed to a person.
See it happen
Inside the loop

- The loop → The tool
- The tool → The system
- The tool → A person
Nobody is watching the loop
The loop calls The tool and decides what to do with the answer, in milliseconds, as many times as its budget allows. Everything a retry policy ever said is still true, and now it is behaviour rather than advice.
Learn more
Why this pattern exists
Everything the reliability lessons say about retrying is still true here, and one thing has changed: nobody is watching. A person who gets a timeout retries, waits, and thinks about whether that was wise. A loop gets the timeout and goes again in milliseconds, and again, until something stops it — so the policy is not advice any more, it is the behaviour of the system. That makes three questions sharper than they were. Is a repeat safe, which is a property of the tool and not of the agent's confidence. Is there anything left in the budget, because a loop with no budget is an outage generator. And is there a person at the end of this, because "stop and ask" is a real answer for an agent in a way it never was for a retry policy in a config file.
An agent is working through a queue of order tasks, and each of these is a moment inside its retry loop. Every row shows the call, what came back, where the budget is, and what the tool does with a repeat. What actually happened at the far end is revealed only after the decision.
- Read a repeat as a property of the tool: an idempotency key, a naturally idempotent call, or neither.
- Keep a loop out of a pause the API asked for, and out of a budget it has already spent.
- Choose between finishing the work, asking the system what happened, handing it to a person, and reporting failure honestly.
The rule this lesson applies: A repeat is safe when the tool says so, not when the agent is confident. An idempotency key makes a repeat a no-op for the same logical operation; a call that sets a value rather than adding one is naturally safe to repeat; anything else may do the work twice, and after a timeout the agent cannot tell whether the first attempt landed. That is the same reasoning as the Timeout lesson, with the loop turning a single bad decision into several. Two things belong to agents specifically. The first is the budget: a loop needs a number of attempts it may spend and a rule for what happens when they are gone, or it will keep going long after a person would have stopped. The second is escalation — an agent can hand the work to a person, which is the right answer exactly when the tool cannot be repeated safely and the effect cannot be undone, and the wrong answer when a safe repeat was sitting there and the budget was untouched. An agent that escalates everything costs more than the process it replaced; an agent that escalates nothing eventually does something nobody can undo.

