AI Progress Live Dashboard

Cross-source synthesis · extrapolation · delta-since-last-look
Synthesis

State of play across the four nodes

live
Generating synthesis…

Extrapolations

  • …

Δ since last look

  • …
Applied capability epoch.ai ↗

SWE-bench Verified — top score

fresh
—
—
scaffold: —
—
Score is agent + scaffold, not raw model capability — the scaffold label keeps a bespoke harness from being misread as model headroom. This is Verified (the 500-sample human-validated subset; Epoch runs 484), not full SWE-bench. Epoch upgraded its scaffold significantly in Feb 2026. Claude Opus 5 released Jul 24 2026 leads at 96.0%; Claude Mythos 5 at 95.5%, Claude Fable 5 at 95.0% (vals.ai, llm-stats.com). Claude Opus 5.5 (released Sep 22 2026) and Claude Fable 5.1/Mythos 5.1 (released Sep 1 2026) have not yet been evaluated on SWE-bench Verified.
Autonomy trajectory metr.org ↗

METR 50% time-horizon — frontier curve + live extrapolation

fresh
— · doubling ~7 mo (2019–25); ~4 mo recent · —
fitted (7-mo doubling) live extrapolation from current anchor 16-h measurement ceiling (flagged May 2026)
Forward markers (extrapolation from current anchor at 7-mo doubling): …
Robustness: METR notes that a 10× absolute-measurement error shifts arrival by ~2 years — slope dominates. Gaps in coverage are not regressions; METR doesn't evaluate every release. Newest points (esp. Claude Mythos Preview) sit at or past the suite's measurement ceiling — wide CI. Claude Fable 5.1 and Mythos 5.1 (released Sep 1 2026) have not been evaluated by METR as of September 2026.

What to reach for this week

fresh
ModelIntel$/Mtoktok/s
AA's Intelligence Index is a composite of ~10 benchmarks (v4.1, updated Jun 2026 to weight agentic workloads, including Terminal-Bench 2.1 and GDPval-AA v2). Methodology shifts; values are relative, not absolute, and reflect benchmark/style biases. Blended $/Mtok shown as 3:1 input:output (a common heuristic, not AA's exact blend). Claude Opus 5 (released Jul 24 2026) leads the frontier at $5/$25 with 60.7 Intelligence Index. Claude Opus 5.5 (released Sep 22 2026) is the latest Opus model at $4/$20 per Mtok; Claude Fable 5.1/Mythos 5.1 (released Sep 1 2026) are $10/$50 per Mtok. None of the three has an Intelligence Index score on this dashboard yet.

Training compute over time

fresh
OWID's AI-compute series is Epoch-derived — same lineage as anything else Epoch-sourced; don't read this as independent corroboration of Epoch's other numbers. Trend lines: 1.5×/yr (1950–2010) → 4.2×/yr (2010–25). Recent frontier runs sit near 10¹¹ petaFLOP (≈10²⁶ FLOP).
Data baked at build time on June 21, 2026 (Cowork artifacts can't reach external APIs). Synthesis runs live in your browser via window.cowork.askClaude. Ask Claude to rebuild this artifact to refresh underlying numbers. SWE-bench Verified leaderboard (Opus 5: 96.0%, Mythos 5: 95.5%, Fable 5: 95.0%), Artificial Analysis Intelligence Index (Opus 5: 60.7, Fable 5: 59.9, GPT-5.6 Sol: 58.9), METR 50% horizons (Claude Mythos Preview 16h cap, Fable/Mythos 5 not yet evaluated), and OWID compute data; Claude Opus 5 released July 24 2026 at $5/$25/Mtok; Claude Opus 5.5 released September 22 2026 at $4/$20/Mtok with 1M context window; Claude Fable 5.1 and Mythos 5.1 released September 1 2026 at $10/$50/Mtok.