AI risk field guide

AI Safety & Existential Risk: A Field Guide

As AI systems approach and exceed human capability, a serious question follows: could this go catastrophically wrong? This hub lays out the case for AI existential risk (x-risk), a way to estimate your own p(doom), the people arguing both sides, and what governance can do — with a deep-dive behind each.

The AI safety landscape

The debate has four moving parts. Get oriented here, then follow each into its full reference.

1. The caseWhy it might go wrong

Misalignment, instrumental convergence, and loss of control — the argument for taking x-risk seriously. AI risk →

2. The oddsEstimating p(doom)

Turn vague dread into an explicit probability you can reason about. p(doom) →

3. The thinkersWho argues what

Yudkowsky and the rationality tradition that framed much of the debate. Yudkowsky →

4. The responseSafety & governance

The ecosystem working on it and how to govern agentic systems. Governance →

The core concepts, in one place

The vocabulary you need to follow any AI safety argument. Each term is one line; the deep dives below expand them.

TermWhat it means
Existential risk (x-risk)A risk that could cause human extinction or permanently curtail humanity's potential.
p(doom)A person's estimated probability that advanced AI leads to catastrophic outcomes for humanity.
AlignmentGetting an AI system to reliably pursue what its designers actually intend, not a proxy of it. See how post-training reshapes likely answers →
AGIArtificial general intelligence: a system matching or exceeding humans across most cognitive tasks.
Instrumental convergenceMost goals imply sub-goals like self-preservation and resource acquisition — dangerous by default.
Orthogonality thesisIntelligence and goals are independent: a highly capable system can pursue almost any objective.
CorrigibilityWhether a system will accept correction or shutdown rather than resist it to protect its goal.
Inner / outer alignmentOuter: the training objective is right. Inner: the learned model actually adopts that objective.
GovernancePolicy, standards, and oversight that constrain how powerful AI is built and deployed.

AI safety & existential-risk cheat sheets

The full deep-dive references behind the landscape and glossary above.

AI existential risk cheatsheet preview
x-riskAGIMitigation

AI Existential Risk Cheatsheet

The core case for AI x-risk: the threat models, why alignment is hard, and the mitigations on the table — the starting reference.

Open guide
AI safety ecosystem hub preview
EcosystemOrgsMap

AI Safety Ecosystem Hub

An interactive map of who works on AI safety: labs, nonprofits, funders, and researchers, and how the field fits together.

Open guide
AI risk timeline cheatsheet preview
TimelineScenariosx-risk

Timeline of a Potential Apocalypse

A scenario timeline of how an AI catastrophe could unfold step by step — the abstract risk made concrete and sequential.

Open guide
p(doom) calculator preview
p(doom)ToolEstimate

p(doom) Calculator

Answer structured questions to turn your intuitions about AI risk into an explicit p(doom) probability you can defend and revise.

Open guide
p(doom) calculator test harness preview
p(doom)TestingMethod

p(doom) Calculator Test Harness

The testing scaffold behind the calculator — how the estimate is validated, for anyone who wants to check the method.

Open guide
Yudkowsky rationality and AI cheatsheet preview
ThinkersRationalitySequences

Yudkowsky: Rationality, AI & The Sequences

The ideas of Eliezer Yudkowsky and the rationality tradition that shaped how the AI-risk debate is framed today.

Open guide
AGI development guide preview
AGIPathsCapabilities

AGI Development Guide

The proposed paths to artificial general intelligence — what would have to be true for AGI, and how close the field is.

Open guide
Governing agentic AI cheatsheet preview
GovernanceAgentsOversight

Governing Agentic AI

Practical controls for agentic systems: reviewed skills, safety invariants, human approval, and observability in real organizations.

Open guide
Interactive map of how LLMs decide what to say
AlignmentBiasModel behavior

How LLMs Decide What to Say

An interactive map of corpus bias, curation bias, and alignment bias — and how post-training reshapes the answers a model is likely to produce.

Open interactive guide