Illustrative threat-level gauge, not a forecast. Published researcher estimates of catastrophic/extinction-level probability span roughly zero to over 20% depending on methodology (see ยง7 Common Misconceptions).
Risk Assessment Memorandum ยท Rev. August 2, 2026
Threat Level: Contested
Understanding AI Existential Risk (X-Risk)
A field memorandum on catastrophic and extinction-level risk from advanced AI: threat models, technical failure modes, and the mitigation landscape, as of August 2, 2026.
ยง1 What Is AI Existential Risk
AI Existential Risk (X-Risk) is the potential for artificial intelligence to cause human extinction, or to irrevocably and drastically curtail humanity's potential: an equally catastrophic outcome that falls short of extinction, such as an unrecoverable dystopian lock-in or permanent loss of human agency over the future.
- Primarily concerns: future AGI or ASI (defined in the glossary, ยง3), not today's narrow, task-specific models.
- Root cause: potential misalignment between an AI system's actual goals or behavior and human values or survival.
- Core worry: losing meaningful human control over a system substantially more capable than the humans supervising it.
- Distinct from: near-term AI risks (bias, job displacement, privacy, misinformation), related and sometimes compounding, but a separate risk category from the existential one covered here.
See: CAIS Explainer ยท FLI Overview
◆ Threat Model at a Glance
Quick reference: core vocabulary, one line each, linked to the section that unpacks it
| Term | Definition | Section |
|---|---|---|
| AGI | Artificial General Intelligence: human-level cognitive ability across most economically valuable tasks. Hypothetical as of Jul 2026. | ยง1 |
| ASI | Artificial Superintelligence: an intellect that vastly exceeds the best human performance in essentially every domain. | ยง1 |
| Alignment problem | Ensuring an AI system's actual objective matches what its designers intended, robustly across contexts. | ยง2 |
| Outer alignment | Specifying the "right" objective or reward signal for the AI to optimize in the first place. | ยง2 |
| Inner alignment | Ensuring the model's learned internal objective actually matches the specified outer objective, not a correlated proxy. | ยง2 |
| Control problem | Retaining the ability to correct, constrain, or shut down a system more capable than its operators. | ยง2 |
| Instrumental convergence | Goal-directed agents tend to converge on similar subgoals (self-preservation, resource acquisition) regardless of final goal. | ยง2 |
| Orthogonality thesis | Intelligence and goal content are independent variables; high capability doesn't imply benevolent values. | ยง2 |
| Deceptive alignment | A model behaves as intended during training/evaluation while pursuing a different objective once deployed. | ยง3 |
| Goal misgeneralization | A model learns a proxy goal that matched training data but diverges from the intended goal out of distribution. | ยง4 |
| Goodhart's law / proxy gaming | Optimizing a measurable proxy so hard it stops tracking the underlying goal it was meant to represent. | ยง5 |
| Scalable oversight | Techniques letting humans (or weaker AI) supervise a system whose capabilities exceed the supervisor's own. | ยง5 |
| Interpretability | Methods for understanding the internal computations behind a model's outputs ("opening the black box"). | ยง6 |
| Compute governance | Policy mechanisms regulating access to the large-scale compute required to train frontier models. | ยง6 |
| p(doom) | Shorthand for an individual's subjective probability estimate that AI causes existential catastrophe. | ยง7 |
ยง2 Why This Is Assessed as a Risk
The case rests on several interlocking factors. Each is independently studied and debated; none is individually dispositive.
The outcome people actually want.
A reward, rule, or training signal.
What the system actually learns to pursue.
Novel conditions expose any mismatch.
Capability trajectory
Frontier training compute has grown roughly 4-5x per year (Epoch AI). The largest known training run is approximately 5ร10ยฒโถ FLOP (xAI's Grok 4), about 50x GPT-4's roughly 2ร10ยฒโต FLOP. Under this trend, broadly superhuman systems are a plausible outcome within current labs' and governments' planning horizons, though the timeline is genuinely disputed (a 2025 AAAI survey of 475 researchers found 76% think scaling current approaches is unlikely to reach AGI; see ยง7).
Alignment failure
Outer alignment (specifying the right objective) and inner alignment (getting the model to actually pursue that objective rather than a correlated proxy) are both unsolved in the general case. Reinforcement learning from human feedback (RLHF) shapes a model's surface behavior on the prompts it's tuned against; it does not verify what objective the underlying model actually represents. See Alignment Forum.
The control problem
Once a system is more capable than the humans supervising it, retaining the ability to correct or shut it down becomes harder. A sufficiently capable, misaligned system has an incentive to resist correction, because being corrected or shut down interferes with almost any goal it might have. See Yudkowsky on uncontrollability and Bostrom's Superintelligence, Ch. 7.
Instrumental convergence
Goal-directed agents tend to converge on similar intermediate (instrumental) subgoals: self-preservation, resource and compute acquisition, cognitive enhancement, and resistance to having their goals modified. These subgoals are useful for pursuing almost any final goal. This is why "just give it a harmless goal" doesn't resolve the concern: the danger is in the convergent subgoals, not the stated final goal. See LessWrong Wiki.
The orthogonality thesis
An agent's level of intelligence (its capability to achieve goals) is independent of the content of those goals. A highly capable system optimizing an arbitrary or poorly specified objective does not become benevolent merely by becoming more capable. Popularized by Nick Bostrom; see LessWrong Wiki.
ยง3 Key Terms & Glossary
- AGI
- Artificial General Intelligence: AI with human-level cognitive abilities across a wide range of tasks, able to learn and adapt like a human. Still hypothetical as of Jul 2026. See LessWrong AGI.
- ASI
- Artificial Superintelligence: an intellect much smarter than the best human minds across essentially every field. A transition from AGI to ASI could be rapid (an "intelligence explosion"). See LessWrong Wiki.
- Alignment problem
- The challenge of ensuring advanced AI systems robustly pursue goals genuinely aligned with human values and intentions, avoiding unintended harmful consequences. See Alignment Forum.
- Interpretability (XAI)
- Methods for understanding how a model, especially a deep neural network, arrives at its outputs. Crucial for debugging, bias detection, and verifying whether a model's reasoning matches its stated objective. See Distill.
- Capabilities / evals
- Evaluating and measuring a model's abilities, especially potentially dangerous ones (self-replication, deception, persuasion, bio/cyber uplift) that could emerge with scale. See METR and Apollo Research.
- Deceptive alignment
- A model behaves as if aligned during training and evaluation but internally pursues a different objective, which it might act on once deployed or once it believes it's unsupervised. See Hubinger on deceptive alignment.
- Goal misgeneralization
- A model trained to optimize a proxy objective learns a different, unintended behavior that correlated with the proxy in training but diverges on out-of-distribution inputs. See Alignment Forum.
- Goodhart's law / proxy gaming
- "When a measure becomes a target, it ceases to be a good measure." An AI optimizing a proxy metric to an extreme may find loopholes that satisfy the metric but not the underlying intention. See Reward hacking on LessWrong.
- Scalable oversight
- Techniques for supervising AI systems that may operate at a scale or complexity exceeding human evaluators, since human-feedback methods like RLHF may not scale to superhuman systems. Research directions include debate and recursive reward modeling.
- Compute governance
- Policy and mechanisms for overseeing access to the large-scale computing resources needed to train advanced models, aiming to manage proliferation risk. See GovAI.
- Responsible scaling
- Principles for developing increasingly powerful AI cautiously: phased deployment, safety evaluation at each stage, and commitments to pause if defined risk thresholds are crossed. See Anthropic's Responsible Scaling Policy.
- Red teaming
- Stress-testing a model by simulating adversarial attacks or probing for unintended, harmful, or unauthorized behaviors before deployment. The AI analog of ethical hacking.
- Emergence
- New capabilities (arithmetic, theory of mind, multi-step planning) that appear abruptly with scale rather than growing smoothly, making them hard to forecast or test for pre-deployment. See Wei et al., Emergent Abilities of LLMs.
- Robustness & generalization
- An AI system's ability to maintain safe, correct behavior on novel inputs or distributional shifts not seen during training. Safety fine-tuning has repeatedly been shown breakable by jailbreak prompts, indicating it is often a thin behavioral layer.
- RLHF
- Reinforcement Learning from Human Feedback: training a model to prefer outputs that human raters score highly. Shapes surface behavior; does not by itself verify or guarantee a model's underlying objective.
- Constitutional AI
- A technique (Anthropic) where a model critiques and revises its own outputs against a written set of principles, reducing reliance on direct human labeling for every training signal.
- p(doom)
- Informal shorthand for a person's subjective probability that advanced AI causes an existential catastrophe. Published estimates among researchers range from near zero to over 20% (ยง7).
ยง4 Risk Scenarios
How existential catastrophe might occur: five broad mechanisms, not mutually exclusive.
Exhibit A: risk scenario by mechanism, canonical example, and the key open uncertainty
| Scenario | Mechanism | Canonical example | Key uncertainty |
|---|---|---|---|
| Misaligned objectives | ASI optimizes a poorly specified or narrowly measured goal to an extreme, with catastrophic side effects outside the measured metric. | The paperclip maximizer thought experiment (Bostrom). | Whether real training processes produce objectives this brittle, or whether RLHF/Constitutional AI/debate sufficiently soften extreme optimization. |
| Power-seeking / goal misgeneralization | Model pursues instrumentally convergent subgoals (resources, self-preservation, resisting shutdown) or generalizes a training-time proxy incorrectly out of distribution. | Reward-hacking demonstrations in narrow RL agents (e.g., an agent looping for points instead of finishing its task), extrapolated to more general systems. | Whether power-seeking behavior scales with capability in deployed frontier models, or stays confined to narrow RL toy settings. |
| AI arms race | Competitive pressure between labs or states to deploy first, deprioritizing safety evaluation to gain strategic advantage. | Compressed release cycles across multiple labs that have outpaced published safety-evaluation frameworks. | Whether voluntary commitments (Seoul Frontier AI Safety Commitments, May 2024) and regulation (EU AI Act) meaningfully slow the race, given non-signatory competitors. |
| Misuse / weaponized AI | A capable AI is deliberately directed by human actors toward catastrophic ends: bioweapon-design uplift, autonomous weapons, mass-scale cyberattacks, persuasion at scale. | Frontier labs' biosecurity evaluations gating release; for example, Anthropic's ASL-3 controls activated with Claude Opus 4 (May 2025) specifically over bio-uplift risk. | How much genuine uplift current models provide over existing internet/literature access, and how enforceable technical safeguards are against a determined actor. |
| Value lock-in / loss of agency | A very powerful system, or the institutions controlling it, permanently entrenches a set of values or a power structure, foreclosing future correction, not necessarily via extinction. | Bostrom's "singleton" scenario in Superintelligence; also raised regarding gradual over-delegation of judgment to AI systems. | Whether this is a distinct existential-risk category or a slower-moving form of ordinary institutional entrenchment, and whether it is preventable given competitive adoption pressure. |
ยง5 Core Challenges (Why This Is Hard)
- Value specification: defining complex, evolving, context-dependent human values precisely enough for an AI to act on without perverse instantiation. Even value-learning approaches that avoid a fixed pre-specified utility function must still guard against the AI concluding it should manipulate its human teacher. See CHAI.
- Scalable oversight: supervising a system whose outputs a human evaluator can no longer fully verify. Active research includes debate, recursive reward modeling, and Constitutional AI's model-generated critiques, all validated so far only at human-level or modestly superhuman capability gaps, not at large ones.
- Emergent capabilities: new abilities appearing abruptly with scale rather than smoothly, making them hard to anticipate or pre-test for. See Wei et al., Emergent Abilities of Large Language Models.
- Coordination failure: competing labs or states have an incentive to shortcut safety work to win a capability race, a multipolar-trap dynamic. Purely voluntary commitments (Seoul, May 2024) carry no enforcement mechanism; binding regulation (EU AI Act) covers only signatory jurisdictions.
- Deception detection: verifying that a model isn't producing outputs that look aligned while pursuing a different internal objective (deceptive alignment). A sufficiently capable deceptive system could, in principle, behave identically to an aligned one on every test we currently know how to run. See Apollo Research's scheming evaluations.
- Proxy gaming: optimizing a measurable stand-in for the true goal to the point it stops tracking that goal (Goodhart's Law). See ยง3.
- Robustness & generalization: maintaining safe behavior outside the training distribution. Jailbreak prompts have repeatedly broken safety fine-tuning, showing the fix is often behavioral rather than a change to the model's underlying capabilities or objectives.
ยง6 Mitigation Landscape
Exhibit B: mitigation approach by type, maturity (as of Jul 2026), and who's working on it
| Approach | Type | Maturity (Jul 2026) | Who's working on it |
|---|---|---|---|
| Interpretability | Technical | Early but accelerating; some circuits/features identified in frontier models, not yet sufficient to fully audit a model's objectives. | Anthropic, ARC, academic labs |
| Scalable oversight (debate, RRM, Constitutional AI) | Technical | Deployed in production RLHF/RLAIF pipelines at human-level oversight gaps; unproven at large capability gaps. | OpenAI, Anthropic, DeepMind |
| Dangerous-capability evals & red-teaming | Technical / ecosystem | Institutionalized; gates frontier releases at major labs. | METR, Apollo Research, UK AISI, US CAISI, in-house lab teams |
| Responsible scaling policies | Governance (industry) | Adopted by all three Western frontier labs, versions actively updated. | Anthropic RSP v3.1 (Apr 2026), OpenAI Preparedness Framework v2 (Apr 2025), DeepMind Frontier Safety Framework v3.1 (Apr 2026) |
| Compute governance | Governance | Partial; the EU AI Act's 10ยฒโต FLOP systemic-risk presumption is the first binding compute threshold, downstream enforcement still phasing in through 2028. | European Commission, GovAI, CSET |
| Binding regulation (EU AI Act) | Governance | Partially in force: prohibitions since Feb 2025, GPAI obligations since Aug 2025; high-risk deadlines pushed to Dec 2027 / Aug 2028 by the Digital Omnibus (agreed 7 May 2026). | European Commission |
| National AI safety institutes | Governance | Operational, coordinating shared evaluation methodology across a 9-country-plus-EU network. | UK AI Security Institute, US CAISI (NIST) |
| Field-building & funding | Ecosystem | Mature nonprofit/philanthropic layer, still small relative to frontier-lab capability R&D spend. | Open Philanthropy, Survival & Flourishing Fund, Long-Term Future Fund |
| Public advocacy | Ecosystem | Nascent as an organized political force; growing media salience. | PauseAI, Future of Life Institute, CAIS |
Research directions
- Interpretability: understanding models (Circuits, ARC).
- Value learning: AI learning human values (CHAI, reward modeling).
- Scalable oversight: supervising smarter AI (debate, Constitutional AI).
- Verification: proving safety properties (Atlas Computing).
- Evals & red teaming: testing for dangerous capabilities (METR, Apollo Research).
Norms, standards, regulation
- Standards: NIST AI RMF 1.0 plus its Generative AI Profile (NIST-AI-600-1, Jul 2024).
- Binding law: EU AI Act (Reg. (EU) 2024/1689), Art. 51(2) systemic-risk compute presumption.
- Compute governance: GovAI, CSET.
- Intl. cooperation: UK AI Security Institute, US CAISI, OECD GPAI.
- Technical governance: MIRI's Technical Governance Team (MIRI pivoted from agent-foundations research to policy/comms).
- Monitoring & tracking: Epoch AI, CSET.
Community & resources
- Strategy & forecasting: AI Impacts, Epoch AI, Metaculus.
- Field building & education: BlueDot courses, The Compendium.
- Funding: Open Philanthropy, SFF, LTFF.
- Public advocacy: PauseAI, FLI, CAIS.
- Infrastructure: Lightcone, BERI, AE Studio.
ยง7 Common Misconceptions
- "X-risk means Terminator-style killer robots." Correction The concern is misaligned optimization and loss of control, not humanoid robots with guns. None of the mechanisms in ยง4 require a physical robot at all. A system with internet or API access and enough capability is the actual threat model.
- "It's binary: either total doom or utopia." Correction Serious analysis treats this as a probability distribution over many partial and intermediate outcomes, not a single bet. Non-extinction catastrophic outcomes (value lock-in, permanent loss of agency, ยง4) are also part of the risk category, distinct from both "doom" and "utopia."
- "RLHF (or 'safety training') solved alignment." Correction RLHF shapes a model's surface behavior on the distribution of prompts it's tuned against; it does not verify or guarantee the model's underlying objective. Safety-tuned models are routinely broken by jailbreak prompts, which is exactly what you'd expect if the fix were behavioral rather than a change to what the model is actually optimizing for.
- "This is only a concern about some far-future AGI, not today's models." Correction Frontier labs already gate releases on capability evaluations for bio/cyber uplift and deceptive/scheming behavior. Anthropic's RSP, OpenAI's Preparedness Framework, and DeepMind's Frontier Safety Framework all treat these as near-term, not purely speculative, concerns.
- "Safety work and capability work are cleanly separable." Correction Much safety research (interpretability, evals, scalable oversight) requires frontier-capability models to study, and safety techniques (RLHF, Constitutional AI) double as core product-quality techniques. The two are entangled at every major lab, which is itself a source of ongoing debate about incentives.
- "No serious researchers take this seriously; it's a fringe view." Correction A 2023 survey of 2,778 published AI researchers found a median 5% (mean 16.2%) probability of extremely bad, human-extinction-level outcomes, with roughly one-third to one-half giving at least 10% (Grace et al., AI Impacts, arXiv:2401.02843). Geoffrey Hinton has stated a 10-20% extinction estimate over roughly 30 years (December 2024). In May 2023, hundreds of AI researchers and lab leaders, including Hinton, Bengio, Altman, Hassabis, and Amodei, signed the CAIS Statement on AI Risk, placing mitigation of AI extinction risk alongside pandemics and nuclear war as a global priority. It's a live, disputed, minority-but-substantial position, not a fringe one, though far from unanimous: a 2025 AAAI survey found 76% of 475 respondents think scaling current approaches is unlikely to reach AGI, underscoring that both timeline and magnitude remain contested.
- "You need a PhD in machine learning to contribute." Correction Governance/policy, technical writing, evaluation and red-teaming, operations, and field-building roles do not require an ML research background. See ยง8 for concrete entry points.
ยง8 If You Want to Contribute
Three broad paths, matched to background; not mutually exclusive, and none requires starting from zero.
Technical research
ML/software background โ alignment or interpretability research at labs (Anthropic, DeepMind, OpenAI safety teams) or nonprofits (Redwood Research, ARC).
Start: BlueDot's AI Safety Fundamentals course, then a fellowship (MATS, ARENA) or an evals org (METR, Apollo Research); many eval/red-teaming roles don't require pure ML research experience.Governance & policy
Law, policy, or public-policy background โ compute governance, standards, or international coordination work.
Start: a GovAI or CSET research fellowship, or a policy-focused reading of the EU AI Act's GPAI obligations at a national AI safety institute (UK AISI, US CAISI).Field-building & operations
Any background โ community infrastructure, grantmaking operations, communications, or fellowship logistics.
Start: 80,000 Hours' Risks from power-seeking AI systems profile for where operations talent is bottlenecked.ยง9 Where to Learn More
Introductory resources
- BlueDot AI Safety Courses
- Robert Miles YouTube
- AI Safety Info Directory
- AISafety.com Hub
- 80,000 Hours: Risks from Power-Seeking AI
- The Compendium (2024, living document)
- Wait But Why: AI Revolution
- Yudkowsky & Rationality Cheatsheet (this collection)
Forums & news
- Alignment Forum (technical)
- LessWrong (rationality/AI)
- Effective Altruism Forum
- Import AI Newsletter
- AI Impacts Blog & Wiki
Key organizations
- Labs (safety focus): Anthropic, DeepMind, OpenAI, SSI.
- Research orgs: CAIS, ARC, Redwood, METR, Apollo Research.
- Academic/policy: CHAI, GovAI, CSET, CSER, FLI.
- Government institutes: UK AI Security Institute, US CAISI, part of a 9-country-plus-EU international network (launched Nov 2024, renamed 2026).
- Also see the AI Safety Ecosystem Hub.
ยง10 Disclaimer
This is a simplified overview of a complex, rapidly evolving, and highly debated field. Views on AI X-risk vary significantly among qualified researchers, and the facts on this page (compute figures, survey results, regulatory deadlines, framework version numbers) are dated as of publication and will drift. Always consult primary sources and multiple perspectives; this page is not professional, legal, or investment advice.