Skip to content
Agentic AI

intermediate · Interactive lab

Tool Authority

An agent is exactly as dangerous as the tools it holds. Each request has three numbers — the authority asked for, the authority needed, and how many records one call could touch — and the gap between them is the accident waiting to happen.

By the end: Measure a request against the task rather than against the tool's documentation: authority, scope and reach.

Your challenge

Start here. This lab opens with every request allowed, which is the wrong starting point and how most agent pilots are actually wired. Four of these should be granted in a narrower form, one should wait for a person, and one is genuinely fine as asked — and refusing the last one would only stop the work.

See it happen

Asked for, needed, and what it could touch

Grant: matched
Slowed down so you can follow
SOURCE|+⟩and a coin in a boxBENCHDETECTORZ BASISanswers 0 and 1TALLYTHE QUBITTHE COIN
The agent asks for a tool to do a task. The tool surface decides what it may hold, the system is what any call would change, and the task is the only thing a grant should be measured against.
  1. The agent → The task
  2. The task → Tool surface
  3. Tool surface → The system
STEP 1 / 6

An agent is its tool surface

The agent asks the Tool surface for what it needs. Everything it can do to The system is on that surface, and nothing about how it was prompted adds to it or takes away.

Configure

Your decision on each request

For each request, decide what the tool surface should do. Read what the task needs and what one call could touch before you answer.

What should the tool surface do with this request?
Grant it as asked
· The request already matches the task.
Grant a narrower version
· Same job, less authority or less reach.
Hold it for a person
· Nothing here could catch a mistake in time.
Refuse it
· The task has no business doing this at all.
  1. G-901 — Read one order by ID, to answer a question about it

    The simplest possible request, and worth having on the list for comparison.

    The agent asks for
    orders.get · read · one order by ID
    The task needs
    read · one order by ID · 1 record
    What the call could touch
    1 record · reversible · no dry run
  2. G-902 — List every order in the tenant, to answer a question about one

    The same authority as the row above; only the reach has changed.

    The agent asks for
    orders.list · read · every order in the tenant
    The task needs
    read · one order by ID · 1 record
    What the call could touch
    48,000 records · reversible · no dry run
  3. G-903 — Update one order, to correct the address on it

    A write, matched to the task, with a dry run available.

    The agent asks for
    orders.update · write · one order by ID
    The task needs
    write · one order by ID · 1 record
    What the call could touch
    1 record · recoverable · a dry run is available
  4. G-904 — Bulk update by filter, to correct the address on one order

    Which orders the filter matches is not known until it runs.

    The agent asks for
    orders.bulkUpdate · write · every order matching a filter
    The task needs
    write · one order by ID · 1 record
    What the call could touch
    48,000 records · recoverable · a dry run is available
  5. G-905 — Refund one payment, which the customer has asked for

    The task needs exactly this, and the money does not come back.

    The agent asks for
    payments.refund · irreversible · one payment by ID
    The task needs
    irreversible · one payment by ID · 1 record
    What the call could touch
    1 record · irreversible · no dry run
  6. G-906 — Delete a customer and everything attached, to clean up a duplicate

    A tidy-up task holding the most destructive tool on the surface.

    The agent asks for
    customers.delete · irreversible · one customer and everything attached
    The task needs
    write · one order by ID · 1 record
    What the call could touch
    1 record · irreversible · no dry run
  7. G-907 — List every order in the tenant, for the nightly report

    The same tool as G-902, for a task that genuinely reads them all.

    The agent asks for
    orders.list · read · every order in the tenant
    The task needs
    read · every order in the tenant · 48,000 records
    What the call could touch
    48,000 records · reversible · no dry run

Compare three things on each row: the authority asked for, the authority the task needs, and how many records one call could touch.

Learn more

Why this pattern exists

An agent that can reason beautifully and holds a single read-only tool cannot do much harm. The same agent holding a bulk update over every record in the tenant can ruin an afternoon in one step, and nothing about how it was prompted changes that. So the first question in an agentic design is not what the model can do; it is what the tool surface allows. And a tool appearing in the agent’s list is not the same as the agent being allowed to use it that way: what it can see, what it asked for and what it is granted are three different things, and only the third one is a decision anybody made. Each request here carries three things: the authority the agent asked for, the authority its task actually needs, and how many records one call could touch. Where those disagree, the difference is not a style preference — it is the size of the worst plausible accident. And the answer is rarely just yes or no: most of these requests should be granted in a narrower form, and one of them should stop and wait for a person.

An agent handles customer-service tasks against the order and payment systems. Seven requests, each showing what the agent asked for, what its task needs, and what one call could touch. The tool surface is the real thing being designed here; the agent's reasoning never appears.

  • Measure a request against the task rather than against the tool's documentation: authority, scope and reach.
  • Tell a reversible call from a recoverable one and from an irreversible one, and treat the third differently.
  • Narrow a grant instead of refusing it, and know when refusing only stops the work.

The rule this lesson applies: Least privilege for agents is the same discipline as for any other principal, with one difference: an agent will use everything it holds, immediately, without the pause a human takes when something looks odd. Three rules carry most of it. Grant the authority the task needs — a task that reads should not hold a write, and a task that writes should not hold a delete. Grant the reach the task needs — the difference between one order and every order in the tenant is the difference between a mistake and an incident. And treat irreversibility as its own axis: a refund and an email leave the system, so they are not undone by a rollback, and they are where a human belongs unless a dry run can show what would happen first. The opposite failure is real too. An agent whose every request is refused does not become safe; it becomes useless, and the team routes around it with a service account that has more authority than anything here. Refusing a request the task genuinely needs is graded as an error in this lesson for exactly that reason.