Agents, tools, and blast radius
The moment a model can take actions, the interesting question stops being accuracy.
An agent is a model that has been given tools (functions it may call) and allowed to loop: decide, act, observe the result, decide again. That loop is genuinely powerful and it changes your risk profile completely, because the model is no longer producing text for a human to evaluate. It is doing things.
There is a second, sharper risk. If an agent reads content it did not author (a web page, an email, a file a user uploaded) that content can contain instructions aimed at the model. This is prompt injection, and the uncomfortable part is that there is no reliable filter for it, because instructions and data arrive through the same channel and look identical.
- Do not grant the tool
The only fully reliable control. Give read access where read access suffices. - Require confirmation
A human approves anything irreversible or outward-facing. - Scope the permissions
The agent’s credentials should be weaker than yours, not equal to them. - Log every action
So you can find out what happened afterwards.
This is the point in the whole curriculum where understanding the rest of it pays off most directly. Blast radius is authorization. Confirmation is a trust boundary. Logging is observability. Irreversibility is the rollback lesson. An agent is not a new discipline. It is every earlier lesson, with the stakes raised.
What to remember
- Agents act, so the question shifts from accuracy to consequence.
- Prompt injection has no reliable filter; limit capability instead.
- Withholding a tool is the only fully dependable control.
Terms in this lesson
Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.