Stop building a specialist agent for every job. Capture process, toolbox and proof in a playbook a general agent can run — and put every correction back into the folder.
4 concepts 4 decision paths Diagrams
Share
The playbook model
The playbook framing below is adapted from Actionable AI's AI Delegation Loop guide at theactionableai.com — it is not mine. What follows is my read for practitioners who keep rebuilding the same agent wiring.
Specialist agents die. Job knowledge does not.
One agent per job carries its own prompt, tools and wiring. When the job changes, the wiring breaks and you find out in production. What a general agent lacks is rarely intelligence — it is your judgment written down as a playbook it can run.
Treat the playbook as the durable asset. The model is replaceable; the process, toolbox and proof for a weekly job are not.
Specialist agents die. Job knowledge does not.
One agent per job carries its own prompt, tools and wiring. When the job changes, the wiring breaks and you find out in production. What a general agent lacks is rarely intelligence — it is your judgment written down as a playbook it can run.
Treat the playbook as the durable asset. The model is replaceable; the process, toolbox and proof for a weekly job are not.
Process, toolbox and proof are different jobs
The process layer is judgment in the order you actually use it: triggers, inputs, steps, decision rules, definition of done and edge cases. The toolbox holds reusable scripts, templates, reference data and examples. The proof layer is a checklist that can fail with evidence outside the draft.
Teams that only write a prompt are still doing quality control themselves. Teams that skip the toolbox rebuild the same template every run — and get a slightly different result each time.
A check that cannot fail is not a check
Self-review — the model reading its own draft and declaring it fine — is worth nothing. A real check produces evidence outside the text: a figure traced to a source file, a live link, a count against a limit, a spelling matched to a reference list, a claim matched to a quote.
Without proof you have not delegated the job. You have delegated the draft and kept the inspection. That is how “90 percent done” becomes your permanent role.
Fix the playbook, not the chat
When output is wrong, diagnose which layer failed: a missing step (process), a rebuilt or stale artifact (toolbox), or a missing check (proof). Make the smallest durable fix in that layer, note it, and re-run from scratch. The chat closes; the folder compounds.
Correcting in the chat trains you to type the same fix every week. Correcting in the folder is how the second run needs less of you than the first.
From chat fixes to a closed loop
Maturity is whether the next run needs less of you than the last one.
Delegation maturity
How durable the job knowledge becomes
Chat corrections
Fixes live in a thread that closes
Resets weekly
Giant prompt
Everything in one growing instruction block
Brittle
Written process
Purpose, steps, decisions and edge cases exist
Repeatable
Process + toolbox
Templates, scripts and examples are frozen
Consistent
Closed loop
Proof checks and notes feed the next run
Compounds
Start with the job you do weekly and would recognise a bad version of instantly. Do not start with the most complicated thing you do.
Delegation decisions
Frequently asked questions
What is the AI delegation loop?
A way of handing a repeated job to a general AI agent by capturing it as a playbook — process, toolbox and proof — and putting every correction back into that playbook. The loop is run, fail, diagnose the layer, fix the folder, re-run. The chat closes; the folder compounds. The framing comes from Actionable AI's AI Delegation Loop guide; this page is my practitioner read of it, not a reprint.
Why not build one specialist agent per job?
Each specialist carries its own prompt, tools and wiring. When the job changes, the wiring breaks and the operator finds out late. Actionable AI's guide points at the same failure Vercel and Anthropic engineers have written about: chaining specialists looked clever, then a general agent plus a folder of the team's own files outperformed the wiring. What was missing was job knowledge, not another agent.
What are the three layers of a playbook?
Process is judgment written down: purpose, triggers, inputs, steps, decision rules, definition of done and edge cases. Toolbox is the reusable scripts, templates, references and examples the process points at. Proof is a set of checks that must pass with evidence outside the draft before anything reaches you.
How do I start in thirty minutes?
Pick the weekly job you recognise a bad version of instantly. Have the agent interview you into a process file, run the job once and freeze whatever came out right into the toolbox, then add five to ten checks that can fail. The second run is where the pattern clicks — stop only after you have seen it.