Build log HIB-26·Claude Code·human-reviewed before merge·audit trail: public
Doc Nº HIB-26-07 Class Engineering Exhibit Agent Claude Code Models Fable 5 → cheaper tier

How This Site
Is Built
the build log

This collection isn't hand-authored one page at a time. It's the output of a governed agentic-AI pipeline: an AI coding agent (Claude Code) working against a version-controlled specification that functions as binding acceptance criteria, with work delegated across model tiers and gated by a self-verification checklist and human diff review before deploy. This page is the exhibit — the architecture, the spec, the governance, and the deliberate engineering choices, with rationale.

Register of controls
Scope160+ referencesone shared spec, one gate
AuthorClaude Codefrontier + cheap tiers, effort = high
DelegationBy costFable specs & reviews; cheap builds
SpecAGENTS.mdone standard, every agent
Review gateChecklist + humannever the agent alone
Audit trail740+ commits, publicevery diff browsable since Apr 2025
Topic signalLive trafficCloudflare views rank what's next
FrameworkBootstrap 5.3.8pinned + SRI on every asset
Build stepNonestandalone static HTML
Theminglight-dark()native, no dual stylesheet

Why This Page Exists

Anyone can prompt a model into one good-looking page. The hard part — the part worth showing — is doing it repeatably, safely, and at scale: dozens of pages held to the same accuracy and accessibility bar, with the rules written down, enforced, and auditable. That's an engineering problem, not a prompting trick.

The interesting artifact isn't any single cheatsheet. It's the pipeline that makes the next one cheap, consistent, and trustworthy.

A repeatable process

Each page runs through the same Generation Protocol — research, outline to three depths, fill to a density floor, then self-verify. The quality comes from the spec, not from a lucky prompt.

Governed, not vibes

The agent has wide latitude to write, but a written acceptance standard and a human reviewer stand between its output and main. Autonomy is bounded by design.

Fully auditable

Every change is a commit. The full history — messages, diffs, per-file revisions — is rendered straight from git on the change-history page. Nothing is hidden.

The Generation Pipeline

A new reference moves through a fixed sequence. The first four steps are the agent's; the last two are the gate. No page reaches main without clearing all six.

Research
Verify first
Check version-sensitive facts against primary sources before writing — close the hallucination gap up front.
Outline ×3
Three depths
Outline fundamentals, working knowledge, edge/advanced — before any prose.
Fill to floor
Density
Populate every section to the density floor: ~20+ substantive entries, no hollow stubs.
Self-verify
Checklist
Run the page against the Testing Checklist — coverage, accuracy, platform, a11y — and fix gaps.
Browser verify
Prove it loads
Serve locally, confirm a clean console, prove the CDN assets actually loaded under SRI.
Review & commit
Human gate
A human reviews, then it ships to main as one commit bundling the page and its preview image.
Delegation by cost

The pipeline spends frontier tokens on judgment and commodity tokens on construction. A frontier model (Claude Fable 5) does the work that needs taste and verification — picking topics against live traffic, authoring the spec, and reviewing the result. Implementation of each page is delegated to a cheaper model, and in batches to parallel workers with disjoint ownership: one worker per page owns its .html and preview image, while shared files (the category map, index.php, batch commits, the deploy) stay with the supervising agent. Comprehensiveness comes from the standard, not from cranking the effort dial — if a draft is thin, the fix is to apply the spec harder, not to spend more tokens.

The loop closes

A nightly GitHub Actions job pulls per-page view counts from Cloudflare GraphQL Analytics into popularity.json, which ranks the gallery and powers the public traffic dashboard. That signal feeds back into the first step: what people actually read informs what gets built next. Ideation → spec → build → deploy → measure → ideation.

The Spec: AGENTS.md as Acceptance Criteria

The pipeline's center of gravity is a single checked-in file, AGENTS.md. It's deliberately cross-agent: CLAUDE.md and .github/copilot-instructions.md both point back to it, so Claude Code, Codex, and Cursor build to one standard instead of three divergent house styles. Its rules are treated as binding acceptance criteria — a page that violates them is not done, regardless of polish.

Excerpt — AGENTS.md § Atomic entry rule
# verbatim, from the checked-in spec
Every entry (concept, command, pattern, term, technique) MUST include:
- A precise one-line definition or statement of purpose.
- At least one concrete example with realistic values
  — never foo/bar when a real value teaches more.
- Where applicable, a gotcha, pitfall, or explicit "when NOT to use this."

## Quantify everything quantifiable.
"Fast" → "~O(log n), sub-ms for n < 10^6."
"Expensive" → the actual price/token figure.
"Large" → the actual cutoff.
What the spec enforcesThe concrete rule
Coverage contractEvery topic covered at three depths — fundamentals, working knowledge, edge/advanced. No section ships hollow.
Atomic entry ruleEvery entry needs a one-line definition, at least one concrete example with realistic values, and a gotcha where applicable. No foo/bar.
Density floorTypically 20+ substantive entries per page; any section under ~3 entries is folded or expanded. Thinness is a defect.
Accuracy gateEvery version, price, limit, or benchmark verified against a primary source — "verify, don't recall." Fabricated specifics fail the gate.
FreshnessVolatile facts dated inline; review status is tracked in refresh-status.json, not as a visible stamp or JSON-LD dateModified on the page — a routine that only bumps a page date without real review is worse than an honest gap.
Structured-data honestyJSON-LD must match the visible content — never describe in schema what isn't on the page.
Platform baselinePinned Bootstrap with SRI, deferred JS, native <details>, light-dark() theming, @layer cascade, WCAG 2.2 AA, Core Web Vitals targets.

The same file also carries the Build & Verify Workflow — concrete steps from past builds (compute SRI from real CDN bytes, serve and confirm Bootstrap actually loaded, generate the preview at 1200×630) that make a page correct on the first pass instead of after a round of review comments.

The Governance Model

"Governed" means specific, checkable things — not a vibe. The agent writes; four mechanisms bound what it can ship and keep the result accountable after the fact.

1A gate that isn't the agent

The Testing Checklist (comprehensiveness, accuracy, platform, accessibility) plus a human reviewer sit between generated output and production. The model proposes; a person reviews the diff and gates the deploy. The gate produces a report, never the final verdict.

2A public, immutable audit trail

Every page is shipped as a version-controlled commit with a descriptive message and an authorship trailer. The change-history page renders the entire git log — commits, diffs, and per-file revisions — read-only, straight from the repository.

3Supply-chain integrity by default

Every CDN <link>/<script> carries a sha384 Subresource Integrity hash plus crossorigin. Hashes are computed from the real CDN bytes and pinned in a table in CLAUDE.md; a tampered or swapped asset simply won't execute.

4The agent in the loop

A GitHub Actions workflow (claude.yml) wires the same agent into the repo: an @claude mention on an issue or pull-request review invokes Claude Code in CI — so the system that authors pages can also participate in their review.

5Delegation with disjoint ownership

When work fans out to parallel workers, each owns exactly one page — its .html and preview image, nothing else. The shared files that could collide — the category map, the gallery index.php, batch commits, the production push — are reserved for the supervising agent. Blast radius is bounded by who is allowed to touch what.

The model, generalized

This site is a concrete instance of a broader pattern: turning ad-hoc AI use into a reviewed, versioned capability with declared safety posture. That pattern is written up in detail in Governing Agentic AI: A Field Guide — the substrate-not-the-agent bet, the Golden Rule, four safety invariants, and a reviewed-skill lifecycle.

Deliberate Tech Choices

Every platform decision here is a trade made on purpose. Expand each for the rationale — what it buys, and what it deliberately gives up.

One documented exception

This exhibit page is the one deliberate departure from the baseline below: it swaps Bootstrap for a bespoke, dependency-light stylesheet — a paper-ledger motif built from CSS custom properties, cascade layers, and scroll-driven animation, no framework — to demonstrate the raw-CSS techniques directly. Every other reference in the collection, including the 150+ cheatsheets this page describes, still runs on the pinned Bootstrap baseline documented in the first item below.

01 Bootstrap 5.3.8, pinned, with Subresource Integrity SRI on every asset

An exact version (5.3.8) is pinned across every page for visual consistency, and Bootstrap 6 is explicitly not targeted until it's stable. Each CDN tag carries an integrity hash and crossorigin; JS loads with defer.

  • Why: a known-good, cached framework gives 150+ pages a shared look without a component build. SRI turns the CDN from a trust dependency into a verified one.
  • Trade-off: bumping a version means recomputing the hash from the real bytes and updating the pin table — friction that's intentional, because a silent version drift is exactly what SRI is there to catch.
02 Native light-dark() theming one token set, both themes

Light and dark are expressed with the CSS light-dark() function plus color-scheme, honoring prefers-color-scheme automatically, with an optional manual override via [data-theme] and a tiny pre-paint script to avoid a flash.

  • Why: one set of design tokens drives both themes — no hand-maintained dual stylesheet to drift out of sync.
  • Trade-off: relies on a modern baseline, which is fine for a 2026 audience and removes a whole class of theming bugs.
03 Native <details> for collapsibles zero-JS accordions

Accordions and task cards use the native <details>/<summary> element (with name for exclusive groups) instead of a JS collapse component.

  • Why: zero JavaScript, accessible by default, and it satisfies the "works without JS" requirement for free — the content is visible and expandable even with scripting off, and expands fully for printing.
  • Trade-off: less animation control than a JS widget — an acceptable price for robustness and accessibility.
04 JSON-LD structured data on every page schema must match content

Each page ships TechArticle JSON-LD whose fields must match the visible content. No FAQPage/HowTo schema is bolted on to chase deprecated rich results.

  • Why: AI answer engines and classic crawlers both read clean structured data; honest schema is the durable lever, not a markup trick.
  • Trade-off: the schema has to be maintained alongside the content — enforced by the rule that structured data can't claim what the page doesn't show.
05 No build step — standalone static HTML one file per page

Each cheatsheet is a single self-contained HTML file with embedded CSS and JS. A small amount of PHP (index.php, sitemap.php, history.php) auto-discovers files and renders the gallery, sitemap, and git history at request time.

  • Why: nothing to compile, bundle, or babysit; any static host serves it, and a page opened directly in a browser still works. Adding a cheatsheet is dropping in one file (plus one line in the category map).
  • Trade-off: some markup repeats across files instead of living in shared components — accepted in exchange for zero build complexity and maximum portability.
06 CSS @layer, container queries, reduced-motion gating predictable cascade

Custom CSS is wrapped in @layer so it sits cleanly above Bootstrap without !important wars; the card grid uses CSS Grid; all motion is gated behind @media (prefers-reduced-motion).

  • Why: predictable cascade, responsive-by-container layout, and motion that respects user preference — modern platform features doing work that used to need JavaScript or specificity hacks.
  • Trade-off: none worth noting at the 2026 baseline; these are strictly-better swaps.

What This Demonstrates

Read as a credential, the relevant skills are visible in the artifact itself, not just claimed — and the numbers are live, not asserted.

References shipped
160+
one spec, one gate
Daily views
~3,500
live traffic dashboard →
Public commits
740+
every diff since Apr 2025

Scale with consistency

160+ references spanning AI safety, software, finance, security, and more — all governed by the same accuracy, accessibility, and metadata bar because the bar is written down and enforced.

Agentic-AI architecture

Designing the spec, the gate, and the audit trail that let an AI agent do real work safely — the substrate, not just the prompt.

Cost-tiered delegation

Spending frontier reasoning where it changes the outcome (topics, spec, review) and cheap tokens on construction — the economics of running a real pipeline, not a one-off demo.

Multi-agent orchestration

Fanning work out to parallel workers with disjoint ownership and a supervising agent holding the shared files — coordination with a bounded blast radius.

Information design

Turning dense, version-sensitive material into scannable, self-contained references that a practitioner can actually work from.