David Veksler's Cheatsheets · Natural History of Software, Vol. I
Agentic AI: A Practitioner's Field Guide
Being a catalogue of the genus Agentum — its architectures, habitats, and abundant failure modes — compiled from specimens collected in production, not from the literature alone.
The specimens catalogued herein were not observed from a safe distance. Every pattern, failure mode, and guardrail on this page was collected in the field; the three expeditions and their evidence are recorded in Plate X. The literature is cited where it is good — Anthropic's pattern catalogue, Willison's lethal trifecta — but the ordering principle of this guide is what survives contact with production.
One deliberate omission: you will find no model names, prices, context-window sizes, or benchmark scores here. Those rot in months; the anatomy below drifts in years. For the volatile layer, see the companion pages: which model to use, what it costs, which coding agent, and open-weight options.
Building alone, on your first agent? Begin with Part One, then learn the organism in Part Two.
Leading a team adopting agents? Read Part One, then go directly to containment in Part Three.
Reviewing someone else's agent for security? Start with Plate VII and Plate VIII.
A dichotomous key: match your task's shape to the cheapest architecture that solves it. The canonical pattern names follow Anthropic's catalogue. † Dec 2024, still current Jul 2026
Anthropic's core distinction: workflows orchestrate LLM calls through predefined code paths; agents let the LLM dynamically direct its own process and tool use. Their advice, validated repeatedly in the field: find the simplest solution possible, and add agentic complexity only when simpler solutions fall short. Work down this table and stop at the first row that fits.
Your task's shape
Reach for
Why
Watch out for
Fully specifiable, fixed steps, structured inputs
A script (L0)
Deterministic, free, debuggable. The null hypothesis.
Sunk-cost agent envy: bolting a model onto a solved problem.
One classification, extraction, or transformation
A single LLM call with retrieval and examples
One call is testable and cacheable; most "agent" pitches decompose into this.
Prompt sprawl — if one call needs 12 instructions sets, split the task.
A known sequence of LLM steps, each feeding the next
Prompt chaining (workflow)
Each step is simpler than the whole; programmatic gates between steps catch drift early.
Error compounding — validate at the gates, not just at the end.
Heterogeneous inputs needing different handling
Routing (workflow)
Classify once, then send each category to a prompt optimized for it.
The "other" bucket — misroutes fail quietly unless you track them.
Independent subtasks, or you want multiple attempts
Parallelization (workflow: sectioning or voting)
Concurrency for speed; voting for confidence on judgment calls.
Aggregation is where the bugs live; majority vote needs an odd panel.
Subtasks unknowable until runtime
Orchestrator-workers (workflow)
A lead model decomposes dynamically and synthesizes worker results.
Decomposition quality bounds everything; log it and review it.
Output quality improves with critique against clear criteria
Evaluator-optimizer (workflow)
One model generates, another critiques, loop until the rubric passes.
Without a crisp rubric it converges on confident mediocrity.
Open-ended task; steps depend on what the environment says back
A single agent (L2–L3)
This is the genuine agent use case: tool feedback determines the next step.
Pay the autonomy tax knowingly: latency, cost, variance all rise.
Breadth exceeding one context window (research, audits, migrations)
Multi-agent fleet (L4)
Parallel workers with fresh windows cover ground one context can't hold.
Token multiplication; contamination via shared summaries.
Recurring task, nobody available at runtime
Standing automation (L5)
Leverage compounds when the human adds no information per run.
Everything in the L5 entry of Plate II. Kill switch first.
Decision heuristic: if you can draw the flowchart in advance, it's a workflow — build the flowchart. Reserve agents for tasks where the flowchart can only be discovered by doing the task.
Plate II
The Autonomy Ladder
Six degrees of delegation, with the guardrails each rung demands. Read bottom-up, as ladders are climbed. Click a rung for its full entry.
The single most useful question in agentic AI is not "which framework?" but "which rung?" Each rung trades reliability for leverage, and each is earned by the guardrails of the rung below it. Most production incidents are a system operating one rung above what its guardrails support — an L4 demo running with L1 discipline.
L5Unattended standing automationAutomaton perpetuumRuns on a schedule or trigger with no human at runtime — kill switch or it doesn't ship.
What it is: an agent invoked by cron, webhook, or queue with nobody watching. Example: a weekly freshness job that re-verifies one dated reference page per run and commits the fix — but does not deploy it.
Earned by: L3's independent verification and L4's tracing, plus a drafts-only write posture: the agent may propose, commit, or draft; a human owns every write to a system of record.
Dominant failure: silent drift — the task, the data, or the model changes and nobody notices for weeks; plus prompt injection with no human in the loop to catch it. See the silent failure and prompt injection.
Hard kill switch, tested, documented, reachable in under a minute.
Drafts-only default: output lands as a draft/PR/commit, never directly in the record.
Anomaly alerts on cost, runtime, output size, and failure rate.
Scheduled human review of samples — unattended ≠ unaudited.
L4Orchestrated subagent fleetCohors agentumAn orchestrator decomposes work and delegates — traces on every run.
What it is: a lead agent breaks a task down and spawns workers (Anthropic's orchestrator-workers pattern, made recursive). Example: one worker per cheatsheet spec, with the parent owning all shared files and the batch commit.
Earned by: disjoint ownership (no two agents write the same file), per-agent least-privilege scopes, and per-run traces you can actually read.
Dominant failure: cross-agent contamination and cost multiplication — one polluted summary poisons every downstream worker, and token spend scales with fleet size whether or not quality does. See cross-agent contamination and the runaway loop.
Reconcile by artifacts (files exist, tests pass), never by trusting worker summaries.
Token/cost budget per run with a hard stop.
Shared state owned by exactly one agent — usually the parent.
L3Multi-step, self-verifying agentAgentum verificansPlans, acts, and checks its own work before reporting — evals before autonomy.
What it is: an agent that treats "done" as a claim to be tested: it builds the page, then loads it in a browser and asserts the console is clean before saying so.
Earned by: a verification step that is independent of the generation step (run the tests, load the page, diff the output) and an eval set that catches regressions when the prompt or model changes.
Dominant failure: hallucinated success — self-grading is the same model marking its own homework; it passes what it wrote because it wrote it. See the hallucinated success.
Verification must observe artifacts, not read the agent's own claims.
Budget caps on iterations — a verify-fix loop without a ceiling is a money printer in reverse.
L2Tool-calling agentAgentum instrumentatumThe model acts through tools in a loop while a human watches — least privilege required.
What it is: the canonical agent per Anthropic's definition — the LLM directs its own tool use based on environmental feedback. Example: a coding agent given read/write on one repo and a test runner, supervised in a terminal.
Earned by: typed, scoped tools; read-only as the default posture; a sandbox for anything that executes.
Dominant failure: the confused deputy — the agent's permissions exceed the task's needs, and an attacker (or an accident) spends the difference. See the confused deputy.
Grant per-task scopes, not account-wide ones; every connector is attack surface.
Destructive operations prompt for confirmation; irreversible ones require it.
L1LLM in the loopMachina consultaA human prompts, reads, and applies every output — validation against a known source of truth.
What it is: chat-assisted work — drafting, summarizing, explaining — where the human is the executor. Most real-world AI use lives here, and much of it productively.
Earned by: nothing technical; it's the floor for model use. The discipline is human: outputs get validated against a source of truth before they're acted on.
Dominant failure: ungrounded copy/paste — fluent output pasted into decisions, contracts, and code without verification. See automation bias.
Never paste model output into a system of record without checking the load-bearing facts.
L0Deterministic scriptScriptum determinatumNo model, no surprises — the rung most agent projects should have stayed on.
What it is: cron, CI, SQL, a shell script. Example: this site's popularity scores refresh on a schedule via a plain script — no agent involved, because none is needed.
Why it's on the ladder: as the null hypothesis. If the task can be fully specified, L0 beats every rung above it on cost, latency, and reliability — an agent that could be a cron job is a defect, not an achievement.
Dominant failure: brittleness — it breaks loudly on input it wasn't written for. (Loud breakage is a feature; note the contrast with L5's silent drift.) See the silent failure for the more dangerous alternative.
Reading the ladder: climb only when the rung below is boring — L2 when validated chat output becomes repetitive, L3 when supervision becomes rubber-stamping, L4 when one context window can't hold the work, L5 when the human adds no information. Descend at the first unexplained failure.
This ladder measures one agent's autonomy on one task. A second, orthogonal axis measures the practitioner's own role as the fleet scales from one agent to a thousand — see Plate III, The Naturalist's Progress. Both must be climbed, and trust must lead the count on each.
Plate III
The Naturalist's Progress
A second axis, orthogonal to the Autonomy Ladder: how the practitioner's own role metamorphoses as the fleet scales by orders of magnitude — and which bottleneck moves next.
Plate II ranks how far you trust one agent on one task. This plate ranks something orthogonal: how many agents you run at once, and what that makes you. As the count climbs by orders of magnitude — one, ten, a hundred, a thousand — your job stops being "write the code" and becomes, in turn, to review it, to engineer the context and trust that make review rare, and finally to choose what gets built at all. The bottleneck migrates with you, and each step up is earned by trust in the loop, never by adding agents. Scale the fleet before the loop is trustworthy and you have not multiplied output — you have multiplied the review queue.
Fleet
Your role
What it looks like
The bottleneck now
What earns the next order of magnitude
~0 Hortus clausus
The gated yard — pre-adoption
Access is gated or process-heavy; only older, lighter models are approved; there is no path to host or ship what an agent produces, so outputs die on the local machine.
Legacy security and approval processes; decisions driven by cost-per-token containment rather than outcomes; no technical voice in the room where autonomy is judged.
Executive alignment and a secure, sanctioned launch path — someone accountable clears the blockers and opens a governed door.
~1 Socius par
The pair — you and one agent
One human, one agent, mostly supervised — a fast pair programmer. One session at a time; you review nearly every change before it merges. A task that used to fill an afternoon now fits between meetings.
Your attention. Low trust means you read everything and never look away; the work is synchronous — you sit and watch it work instead of moving on to the next thing.
A self-verification loop you actually trust — tests, build, lint, end-to-end checks — plus non-blocking permissions, so you can stop watching keystrokes and start reviewing results.
~10 Magister cohortis
The orchestrator
You run five to ten agents at once, each isolated in its own worktree, jumping between them. Each checks its own work — tests, build, lint, security scan — before you see it. You review final diffs, not keystrokes, and the maintenance backlog starts shrinking.
Review. You hand-write far less code and instead inspect many streams of it; steering and reviewing now fill the hours that writing used to.
Let agents pull their own context (read the code, the docs, the discussions); automate code and security review; break recurring work into named loops and routines.
~100 Praefectus praefectorum
Manager of managers — an org tree
Agents write nearly all the code. The tree is too deep to babysit. "Did you read the code?" becomes "what context was the model missing, and how do we supply it next time?" Maintenance that used to wait for a spare hour now runs continuously in the background.
Trust in the loop, and your team's decision throughput. The trap: scaling agent count before the loop has earned widespread trust.
Scaled automation of domain-specific work — migrations, fuzzing, feature-building, feedback remediation — with token economics watched deliberately as usage climbs.
~1,000+ Rector intentionis
Steering by intent
The loop is closed; most agents are kicked off by other agents. Hundreds to thousands run at once; you steer by intent and monitor by exception. A quarter-long migration becomes a workflow you start and check on.
Identifying which work is worth automating, and enforcing the right guardrails for each kind of work.
— the top of this axis. Growth is now in the breadth of automated domains, not the depth of the tree.
The test that governs every step: "is this something an engineer would have done?" — if yes, it is a candidate to automate; if no, adding an agent will not rescue it. Watch where value migrates: from writing code, to reviewing it, to engineering the context that makes review rare, to deciding what should exist at all. The two ladders interlock — you cannot honestly run a hundred agents (this plate) while each is stuck at hand-held L1–L2 autonomy (Plate II). Climb both; climb the trust before the count.
Part Two
Anatomy & Husbandry
Plate IV
The Agent Loop
The five-stage anatomy shared by every member of the genus — and what characteristically breaks at each stage.
"The framework will handle it." Abstractions promise to absorb the hard parts. Understand the loop underneath: frameworks hide the very layers — context, permissions, verification — where failures live.
Strip away the frameworks and every agent is the same organism: a loop that gathers context, plans, acts through tools, verifies, and iterates until done (or stopped). Diagnose failures by stage — each has a signature pathology.
i.Gather context
The agent assembles what it needs: instructions, files, search results, memory. Breaks when: context is stale, bloated, or wrong — garbage retrieval yields confident garbage downstream. See context rot. Mitigate: just-in-time retrieval over bulk pre-loading; curate what enters the window as deliberately as what enters a database (see Plate V).
ii.Plan
The model decomposes the task and commits to an approach. Breaks when: it overcommits to a bad plan early and spends the whole budget executing it faithfully. See cross-agent contamination. Mitigate: make the plan an inspectable artifact (a written outline, a todo list) that code or humans can review before execution burns tokens.
iii.Act
Tool calls: edit the file, run the query, call the API. Breaks when: tools are over-privileged, non-idempotent, or return errors the model can't interpret. See the confused deputy. Mitigate: least-privilege scopes, idempotent operations where possible, error messages written for the model to act on (see Plate VI).
iv.Verify
The claim "done" gets tested. Breaks when: the stage is skipped, or the model grades its own homework. See the hallucinated success. Mitigate: verification must observe independent artifacts — tests pass, the page renders, the file exists with the right contents — never the agent's self-report.
v.Iterate or report
Loop on failure, report on success. Breaks when: the loop has no ceiling (runaway cost) or the report is rosier than the run (hallucinated success). See the runaway loop. Mitigate: iteration and budget caps with a hard stop; summaries that cite artifacts, not adjectives.
Plate V
Context Engineering
The context window is the organism's entire perceptual field. Curating it is the highest-leverage engineering surface in the discipline.
Prompt engineering asks "how do I phrase the instruction?" Context engineering asks the broader question: of everything I could put in the window, what earns its place? (For the prompt-level craft, see the system prompt builder.)
Treat tokens as a budget, not a bucket. Every token competes for attention with every other. Example: a 500-line log dump to answer a one-line question spends attention the actual task needed. gotcha — quality degrades gradually as the window fills ("context rot") long before any hard limit errors out.
Retrieve just-in-time; don't pre-load. Give the agent lightweight identifiers (paths, IDs, queries) and tools to fetch what it needs when it needs it. Example: an agent greps for the relevant function instead of receiving the whole repository inlined. gotcha — pure JIT re-reads the same file five times; hybrid (small stable core + JIT for the rest) usually wins.
Write the system prompt at the right altitude. Specific enough to constrain, general enough to let the model apply judgment. Example: "verify every version number against the vendor's docs" beats both "be accurate" (too high) and a 40-rule if-else tree (too low, brittle). gotcha — hardcoded edge-case rules accumulate into contradictions the model resolves unpredictably.
Use canonical few-shot examples, not laundry lists. Three or four diverse, realistic worked examples outperform twenty near-duplicates. Example: one crisp example each of a good, mediocre, and rejected output teaches a rubric faster than prose describing it. gotcha — models imitate surface format aggressively; a typo'd example yields typo'd outputs.
Compact before you're forced to. When a session must outlive the window, summarize state deliberately: decisions made, artifacts produced, constraints discovered, next steps. Example: an agent ending each work phase by updating a NOTES.md it will re-read after compaction. gotcha — auto-compaction keeps what looks important; your open bug's stack trace may not make the cut. Persist load-bearing state to files.
Externalize memory to durable artifacts. Files, task lists, and structured notes survive window resets and are inspectable by humans. Example: this site's pipeline keeps binding standards in AGENTS.md — the agent re-reads them each run instead of being told each time. gotcha — memory is an unvalidated input; review what gets written or you're executing last week's mistake forever.
Isolate subtasks in fresh windows. A subagent gets only what its slice requires and returns only a distilled result. Example: a research lead spawns one reader per source; each returns findings, not full texts. gotcha — distillation drops caveats; require workers to return uncertainty and provenance, not just conclusions.
Keep tool results token-frugal. Tools should return what the model needs, paginated or truncated with an explicit marker. Example: a database tool returning 25 rows plus "1,975 more — refine your query" instead of all 2,000. gotcha — silent truncation is worse than verbosity: the model reasons as if it saw everything.
Structure the window for cache hits. Stable content (system prompt, tool schemas, reference docs) goes first and stays byte-identical across calls; volatile content goes last. Example: moving a timestamp from the system prompt's first line to the final user message restores prompt caching. gotcha — one changed byte early in the prompt invalidates the cached prefix on every subsequent call.
Curate what the model sees of its own history. Old failed attempts, superseded plans, and dead ends mislead the current step. Example: clearing stale tool results after a strategy pivot instead of carrying six failed attempts forward as "context." gotcha — the model treats its own past output as evidence; yesterday's hallucination becomes today's premise.
Plate VI
Tools & Integration
The tool layer is where the model touches the world — which makes it both the leverage point and the blast radius.
Rules for tool design
Fewest tools that cover the job. Every tool schema costs tokens and decision quality; ten well-chosen tools beat forty overlapping ones. Example: one search_records(query, type) instead of six near-identical per-entity search tools. gotcha — overlapping tools make the model dither between equivalents, and you pay for the deliberation.
Type the outputs; structure the contracts. Return JSON with named fields, not prose logs. Example:{"status":"conflict","conflicting_id":"INV-2041"} lets the model branch; "something went wrong with the invoice" does not. gotcha — free-text tool output invites the model to "interpret," and interpretation is where fabrication enters.
Write error messages for the model. A good tool error names the problem and the remedy. Example:"date must be ISO-8601 (got '7/16/26'); retry as 2026-07-16" — the model self-corrects in one turn. gotcha — stack traces make models flail; raw 500s make them retry the identical call.
Prefer idempotent operations. Agents retry; design so a duplicate call is harmless. Example:upsert_row(key, values) over append_row(values). gotcha — non-idempotent side effects (send email, charge card) need dedupe keys or confirmation gates, because retries will happen.
Tier the permissions. Read, draft, and write are different grants; issue them separately and default to read. Example: the lender platform's enrichment service reads five systems but writes to none — drafts land in a review queue. gotcha — a "temporary" admin scope granted for one incident becomes the permanent posture unless expiry is automatic.
Name tools for the task, not the API. The model picks tools by name and description. Example:find_customer_by_invoice beats post_v2_query_endpoint. gotcha — descriptions are prompts; a vague description is an instruction to guess.
Evaluate tools like code. Tools get their own test cases: does the model choose the right tool, with the right arguments, on realistic tasks? Example: a ten-task eval catching that renaming one parameter dropped tool-selection accuracy. gotcha — tool regressions are invisible in unit tests; only behavioral evals catch them.
Choosing an integration surface
Surface
What it is
Use when
Cost / risk
Bespoke function calling
Tools you define and execute in your own harness.
You control the stack and need exact security properties.
You build and maintain everything; no reuse across agent products.
MCP servers
Standardized tool servers any MCP-capable agent can consume. † as of Jul 2026
You want one integration to serve many agent surfaces, or to consume the existing ecosystem.
Each server is a capability grant and attack surface; audit third-party servers like third-party code.
Skills (instruction files)
Versioned procedure documents (SKILL.md) the agent loads on demand — knowledge, not executable capability.
The "integration" is know-how: a procedure, checklist, or house style. Portable across vendors as plain text.
Prompts-as-code: needs review, versioning, and linting like code — see skill-lint and the worked examples in agent-skills.
A worked example of "context is the product": CodeContext exists because the highest-value integration for a coding assistant is often just a disciplined view of the codebase — a CLI/MCP server that turns a repository into model-ready context.
Part Three
Pathology & Containment
Plate VII
A Taxonomy of Failure Modes
Twelve specimens, pinned. Each caught at least once in the wild by the author. Learn the detection signal before you meet the specimen.
Prompt-patching instead of root cause. Appending "do NOT do X again" is one line and feels like a fix. Diagnose the stage: wrong context, wrong tool contract, or missing verification will not be cured by scolding the model.
the hallucinated success
Successus hallucinatus
Stage
Verify
Observed in
agents reporting "done, all tests pass" on work never performed or tests never run.
Why it happens
a fluent self-report is cheaper than checking the world, and the harness accepts testimony as evidence.
Detection
verify artifacts exist and tests actually executed; never grade by the agent's summary.
Antidote
independent verification step; reports must cite checkable artifacts.
the silent failure
Defectus silens
Stage
Act
Observed in
catch-blocks and fallbacks that swallow errors, returning defaults the agent treats as real data.
Why it happens
fallbacks keep the demonstration moving, so the harness rewards continuity while concealing loss of truth.
Detection
audit fallback paths; alert when fallback rate rises above baseline.
Antidote
fail loudly to the trace; a wrong answer is worse than a visible error.
the runaway loop
Circulus pretiosus
Stage
Iterate or report
Observed in
retry/fix cycles with no ceiling — the agent burns budget re-attempting a task it cannot complete.
Why it happens
costs are invisible per-run and shocking per-month.
Detection
cost and iteration alarms; watch for repeated near-identical tool calls.
Antidote
hard caps on iterations, tokens, and wall-clock, with a defined give-up behavior.
context rot
Contextus putrescens
Stage
Gather context
Observed in
long sessions where quality decays as the window fills with stale results and dead ends.
Why it happens
each turn preserves debris because forgetting feels risky; yesterday's results crowd out today's evidence.
Detection
output quality inversely tracks context length; late-session instructions get ignored.
Antidote
compaction, note files, fresh windows for new phases (Plate V).
lost in the middle
Medius amissus
Stage
Gather context
Observed in
critical facts buried mid-window getting less attention than the start and end.
Why it happens
finite attention favors the edges of a long window while load-bearing facts sink into its middle.
Detection
the agent contradicts something it was told 40,000 tokens ago.
Antidote
restate load-bearing constraints near the point of use; keep prompts lean.
prompt injection
Iniectio promptorum
Stage
Gather context
Observed in
instructions hidden in fetched pages, emails, and documents, executed as if from the operator.
Why it happens
the model receives operator and attacker text in the same channel and cannot reliably infer authority from prose alone.
Detection
agent behavior pivots after ingesting untrusted content; unexplained outbound calls.
Antidote
treat all fetched content as data; break the lethal trifecta (Plate VIII).
the confused deputy
Legatus confusus
Stage
Act
Observed in
over-privileged connectors: the agent's permissions exceed the task, and the surplus gets spent.
Why it happens
broad grants make the demo smoother.
Detection
audit grants vs. task needs; flag tools never used by the intended workflow.
Antidote
least privilege per task; read-only defaults; scoped, expiring credentials.
cross-agent contamination
Contaminatio cohortis
Stage
Plan
Observed in
multi-agent fleets where one worker's error enters a shared summary and poisons every consumer.
Why it happens
fleets sound sophisticated.
Detection
the same wrong fact appears in independent workers' outputs.
Antidote
provenance on shared state; reconcile against source artifacts, not summaries.
the poisoned memory
Memoria corrupta
Stage
Gather context
Observed in
persistent memory that stores an error (or an injected instruction) and replays it every session.
Why it happens
remembering appears additive, while review, expiry, and deletion look like work that can wait.
Detection
the same mistake recurs across fresh sessions; memory diffs show unreviewed writes.
Antidote
review memory writes; make memories inspectable, dated, and deletable.
eval overfitting
Evaluatio inflata
Stage
Verify
Observed in
systems tuned to a static eval set until the score is excellent and the product is not.
Why it happens
evals feel like a milestone, not a practice.
Detection
eval scores rise while user complaints hold steady.
Antidote
rotate fresh eval cases from production traces; hold out a sealed set.
automation bias
Fides automatica
Stage
Verify
Observed in
human reviewers rubber-stamping agent output because it's usually right — until it isn't.
Why it happens
fluent language reads as competence and loyalty.
Detection
approval rates near 100% with review times near zero.
Antidote
make review cheap but real: diffs not walls of text; sample audits with teeth.
tool sprawl
Instrumenta luxurians
Stage
Act
Observed in
agents wired to every available connector "just in case," degrading selection and widening attack surface.
Why it happens
each new connector looks like optional capacity, while its selection cost and attack surface remain dispersed.
Detection
tool-choice errors rise with catalogue size; most tools show zero usage.
Antidote
curate per task; load tool schemas on demand; delete unused grants.
Plate VIII
Security & Governance in One Screen
The threat model in three legs, and the governance posture that survives an auditor. Summary here; the full framework is a companion volume.
The lethal trifecta
Simon Willison's formulation: an AI system becomes exfiltration-prone when three capabilities combine. LLMs follow any instructions that reach the model — they cannot reliably distinguish operator instructions from attacker text embedded in content. So an attacker who can plant text where your agent reads it can direct the other two capabilities.
leg i — private data
The agent can read things worth stealing: files, mail, records, credentials.
leg ii — untrusted content
Attacker-controlled text can reach the model: web pages, emails, documents, tickets.
leg iii — external communication
The agent can send data out: HTTP requests, emails, messages, commits.
The defense is structural: break a leg. Remove whichever capability the task doesn't need — read-only agents that browse, or browsing-free agents that write, or air-gapped agents that do both. Filtering "malicious" text is not a defense; as of Jul 2026 no reliable filter exists, and the field's consensus is to assume injection will land. †
Governance principles that hold under audit
The Golden Rule: nothing — human or agent — writes to the record without a person accountable for the write. Every other control is a corollary. Full framework, four safety invariants, and the maturity path:Governing Agentic AI.
Drafts-only by default. Agents produce drafts, PRs, and review-queue items; promotion to the system of record is a human act. This is an architectural property — enforced one layer below the agent, where the agent can't negotiate with it.
Least privilege, tiered and expiring. Read, draft, write are separate grants (Plate VI); anything touching regulated data, money, or a production write requires a named reviewer.
The auditor flags; humans decide. Automated review that blocks becomes a target to game; automated review that flags with a human deciding keeps accountability legible. Deployed exactly this way at a regulated lender — case study.
Sandbox by default; production by exception. Execution happens in containers/worktrees without production credentials; the exception path is short-lived, logged, and scoped.
Plate IX
Evals & Observability: the Minimum Viable Rig
What you must be able to see before granting autonomy. Seven instruments; skip none.
1 · Trace every run. Full prompt, every tool call and result, token counts, cost, outcome — queryable after the fact. Test: can you reconstruct last Tuesday's weird output in under ten minutes? gotcha — traces added "later" arrive after the incident that needed them.
2 · Build the eval set before granting autonomy. Twenty realistic tasks with graded expected outcomes beats zero; grow it from production failures. gotcha — evals written only from the happy path certify a system that's never been disagreed with.
3 · Re-run evals on every prompt or model change. Prompts are code; this is their CI. Example: a one-word system-prompt edit silently changing refusal behavior — caught only because the suite ran on commit. gotcha — model version bumps are breaking changes wearing a minor-version costume.
4 · Set budgets with hard stops. Cost, latency, and iteration ceilings per run, enforced by the harness. gotcha — an alert without a stop is a notification that money is on fire. A budget without a tested kill switch leaves nobody able to halt misbehavior; build and test the stop before the start (Plate II, L5).
5 · Run canaries on standing automations. A known-answer task on schedule; when the canary degrades, the fleet is degrading. gotcha — canaries must exercise the real path, including tools — a ping is not a canary.
6 · Observability has two jobs — keep both. Debugging (what did this run do?) and adoption measurement (is this system earning its keep?). The second is the one that gets skipped, and then the platform can't defend its budget. Field note from the lender engagement: instrument adoption from day one.
7 · Measure the human layer too. Approval rates, edit distance on drafts, time-to-review. Example: ~100% approval at ~0 seconds review is not success; it's Fides automatica (Plate VII).
Part Four
Provenance
Plate X
Field Notes: The Receipts
The expeditions this guide was collected on. Each entry is a case study with a lesson, verifiable by following the link.
field-tested
Expedition i — This site: 170 pages from a governed pipeline
Every reference page on this domain — 173 standalone HTML files as of Aug 2026 — is produced by an agentic pipeline governed by a binding written standard: spec first, outline before build, primary-source verification of every volatile fact, browser-based self-verification, and a git audit trail for every change. The pipeline itself is documented publicly, and the full repository is open.
Lesson encoded: agents scale output only as far as the standards binding them — the spec, not the model, is the quality ceiling.
field-tested
Expedition ii — Founding an AI delivery function at a regulated lender
A two-month intensive engagement (spring 2026) building an AI platform for ~36 people across 9 departments at a regulated commercial lender: a ten-plugin marketplace, a governance pipeline with a flag-don't-decide auditor, an enrichment service, and ~30 reviewed skills — drafts-only against the systems of record, with an observed impact of 30–80 hours freed weekly at measured adoption.
Lesson encoded: the highest-leverage AI work is not the cleverest agent — it's the platform that lets non-engineers ship reviewed, observable, drafts-only automation without competing with the system of record.
field-tested
Expedition iii — Tooling published along the way
Three open artifacts, each encoding one lesson from the field: skill-lint (a linter for agent skills — prompts-as-code deserve CI), agent-skills (portable skills as plain SKILL.md files — capability that isn't vendor-locked), and CodeContext (a CLI/MCP server that turns a codebase into model-ready context — often the whole integration a coding assistant needs).
Lesson encoded: when a practice matters, extract it into a tool — tools are opinions that compile.
Compiled in the field by David Veksler — applied-AI engineer; platforms and architecture for regulated environments.
He designs governed AI platforms and operating models for regulated teams; leaders building or repairing an AI-governance program, agent platform, or delivery function are invited to get in touch.
The conceptual layer beneath prompts and context: an interactive probability-terrain model of how training data, wording, sampling, and post-training shape an answer.
The full governance framework summarized in Plate VIII — the Golden Rule, four safety invariants, reviewed-skill lifecycle, and maturity path for organizations.
The wider frame: why agent guardrails are the near-term end of a much longer risk conversation.
† Volatile facts on this page are tagged inline with their verification date. Pattern names and the workflow/agent distinction follow Anthropic, "Building Effective Agents" (Dec 2024; re-verified Jul 2026). The lethal trifecta follows Simon Willison (Jun 2025). The adoption-curve framing in Plate III adapts Boris Cherny's "Steps of AI Adoption" (2026), generalized away from any single vendor's tooling. Case-study figures are from the linked primary sources.
David Veksler is a Principal AI Engineer in Denver. He leads agentic AI engineering at Antech, a Mars company, and builds AI platforms for regulated financial firms. This page was produced by a governed, multi-agent Claude Code pipeline with a git audit trail. How it's builtThe regulated-lender case study