Agents don't fail on the model. They fail on the tooling.
A practical playbook for the person who has been handed “we should do something with agents” and is expected to come back with a plan. Fifteen chapters on tool design, connectors, boundaries, failure modes, evaluation and a ninety-day rollout - written for a real company, not a demo.
Tooling Playbook
Three things that decide whether an agent survives contact with users.
Most "the model is dumb" incidents are a tool that returned an unhelpful error or a description that never said when not to use it. Chapter 3 is ten rules; chapter 4 shows the rewrite in full.
Four sentences - whose identity it acts as, what it may read and write, what needs a human, where the record lives. An afternoon up front instead of months of retrofitting.
Thirty real tasks with written outcomes beats any benchmark. Four metrics tell you when to widen the boundary and when something has quietly broken.
Fifteen chapters, no filler.
Every chapter is one page: a table you can act on, a worked example, or a checklist. It assumes you can read a JSON snippet and nothing more.
The one distinction - the model picks the next step - and everything it costs you.
Model, tools, context, boundary. Four parts, four characteristic failures.
Ten rules for tool definitions a model can actually use.
The same capability before and after, in full, with the error messages.
What the protocol fixed, and the five questions it leaves to you.
What to buy, what to build, and how to tell which is which.
Why one model for every step quietly caps the quality of all of them.
A decision table, plus the pattern that pays first in most companies.
Loops, confident wrongness, tool sprawl, permission creep - and the guardrails.
The trifecta, and the mitigations that work at the tooling layer.
Approval gates by action class, and what makes an approval real.
Thirty golden tasks, trace review, and four metrics worth a dashboard.
Six windows, and the three ways this goes wrong.
Who asks, what they ask, and what answers it.
Twenty boxes before an agent touches production.
“An agent is a system where the model chooses the next step, and the number of steps is not known when you press go. Everything else - tools, memory, planning, multi-agent choreography - is implementation detail layered on top of that one property.”
Every agent your company runs, owned, scoped and logged.
The playbook works on any stack. If you'd rather not assemble the platform layer yourself, StickyPrompts brings every model, connector, approval gate and run log into one governed workspace.