AI in the stackLesson 4 of 49 min

Agents, tools, and blast radius

The moment a model can take actions, the interesting question stops being accuracy.

An agent is a model that has been given tools (functions it may call) and allowed to loop: decide, act, observe the result, decide again. That loop is genuinely powerful and it changes your risk profile completely, because the model is no longer producing text for a human to evaluate. It is doing things.

There is a second, sharper risk. If an agent reads content it did not author (a web page, an email, a file a user uploaded) that content can contain instructions aimed at the model. This is prompt injection, and the uncomfortable part is that there is no reliable filter for it, because instructions and data arrive through the same channel and look identical.

Containment, in decreasing order of effectiveness
  1. Do not grant the tool
    The only fully reliable control. Give read access where read access suffices.
  2. Require confirmation
    A human approves anything irreversible or outward-facing.
  3. Scope the permissions
    The agent’s credentials should be weaker than yours, not equal to them.
  4. Log every action
    So you can find out what happened afterwards.

This is the point in the whole curriculum where understanding the rest of it pays off most directly. Blast radius is authorization. Confirmation is a trust boundary. Logging is observability. Irreversibility is the rollback lesson. An agent is not a new discipline. It is every earlier lesson, with the stakes raised.

What to remember

  • Agents act, so the question shifts from accuracy to consequence.
  • Prompt injection has no reliable filter; limit capability instead.
  • Withholding a tool is the only fully dependable control.

Terms in this lesson

Field notes

Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.

The citation that did not exist

A support bot answered a policy question with a confident reference to section 4.2 of a document that has no section 4.2. Both the answer and the citation were generated the same way, which is exactly why one looked as trustworthy as the other.

Grounding matters

A two-cent feature and a four-figure bill

A summarisation feature cost about two cents per use. It was invisible in testing. At launch volume it was several thousand dollars a month, discovered on an invoice rather than in a design review.

Cost per call times realistic volume

resolved in 900ms · region iad1

Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.