Where each section goes

A & B under the one-course + offload plan · draft mapping, editable in generate_compare.py ← Course A
→ Combined — Kept in the merged A+B course → Strategy+ — Offloaded to AGI Strategy (new overview unit) → Technical+ — Offloaded to the Technical course ✂ dropped — Cut
Browse the editable course mockups: Combined (A+B) · AGI Strategy · Technical
Course A
The Alignment Problem
10 Combined 9 Strategy+ 4 Technical+
Unit 1 What is AI Alignment?16
What is intelligence?Strategy+
A collection of definitions of intelligence
Shane Legg & Marcus Hutter · 2m
🎬
The Power of Intelligence
Rational Animations · 7m
The Design Space of Minds-In-General
Eliezer Yudkowsky · 3m
Clarifying and predicting AGI
Richard Ngo · 5m
Failure by defaultStrategy+
🎬
What happens if AI just keeps getting smarter?
Rational Animations × ControlAI · 14m
🎬
The real risk of AI
Siliconversations · 6m
But why would the AI kill us?
Nate Soares · 3m
What failure looks like
Paul Christiano · 11m
The date of AI Takeover is not the day the AI takes over
Daniel Kokotajlo · 3m
Further readingfurther reading
🎮
Frank Lantz · ~25m
The Alignment ProblemCombined U1
Clarifying "AI alignment"
Paul Christiano · 6m
Vinge's Principle
Eliezer Yudkowsky · 6m
Cognitive Uncontainability
Eliezer Yudkowsky · 5m
Further readingfurther reading
Eliezer Yudkowsky · 15m
Peli Grietzer & Tushita Jha · 7m
Unit 2 Thinking about powerful systems10
OptimisationCombined U2
Optimality is the tiger, and agents are its teeth
Veedrac · 20m
The ground of optimization
Alex Flint · ~18m excerpt
Further readingfurther reading
Eliezer Yudkowsky · 7m
The difference between values and capabilitiesCombined U2
Sorting Pebbles Into Correct Heaps
Eliezer Yudkowsky · 6m
🎬
Intelligence and Stupidity: The Orthogonality Thesis
Robert Miles · 13m
Terminal Values and Instrumental Values
Eliezer Yudkowsky · 13m
Predictable behaviourCombined U2
🎬
Why Would AI Want to do Bad Things? Instrumental Convergence
Robert Miles · 11m
What's Up With Confusingly Pervasive Goal-Directedness?
Raymond Arnold · 5m
Further readingfurther reading
Nick Bostrom · PDF, ~38m e
Unit 3 How goals go wrong18
The hidden complexity of wishesCombined U3
The Hidden Complexity of Wishes
Eliezer Yudkowsky · 10m
You Are Not Measuring What You Think You Are Measuring
John Wentworth · 10m
Further readingfurther reading
John Wentworth · 14m
Story time✂ dropped
Scott Alexander · ~5m
Eliezer Yudkowsky · 12m
Tomás Bjartur · 19m
Tomás Bjartur · 18m
Goodhart's lawCombined U3
Eliezer Yudkowsky · 8m
Too much efficiency makes everything worse: the strong version of Goodhart's law
Jascha Sohl-Dickstein · ~20m
Further readingfurther reading
Goodhart Taxonomy
Scott Garrabrant & David Manheim · ~12m
WireheadingTechnical+
🎬
9 Examples of Specification Gaming
Robert Miles · 10m
Nearest unblocked strategy
Eliezer Yudkowsky · 6m
Edge Instantiation
Eliezer Yudkowsky · 5m
Further readingfurther reading
Inner alignmentCombined U3
🎬
The OTHER AI Alignment Problem: Mesa-Optimizers and Inner Alignment
Robert Miles · 23m
Issa Rice · ~5m
Unit 4 Core challenges of the alignment problem14
DeceptionTechnical+
🎬
Deceptive Misaligned Mesa-Optimisers? It's More Likely Than You Think
Robert Miles · 10m
Deep Deceptiveness
Nate Soares · 18m
Why we should expect ruthless sociopath ASI
Steven Byrnes · 10m
Hidden reasoningTechnical+
Oversight Misses 100% of Thoughts The AI Does Not Think
John Wentworth · 2m
Steganography in Chain of Thought Reasoning
A Ray · 7m
Self-confirming predictions can be arbitrarily bad
Stuart Armstrong · 6m
The Parable of Predict-O-Matic
Abram Demski · 18m
Further readingfurther reading
The difficulties of alignmentCombined U4
Wikipedia · ~10m
🎬
Further readingfurther reading
Eliezer Yudkowsky · ~4m
Eliezer Yudkowsky · ~36m
Unit 5 Surrounding environment14
The situationStrategy+
Robustness to Scale
Scott Garrabrant · 2m
"Sharp Left Turn" discourse: An opinionated review
Steven Byrnes · ~12m excerpt
Irretrievability; or, Murphy's Curse of Oneshotness upon ASI
Eliezer Yudkowsky · ~26m
AI as a science, and three obstacles to alignment strategies
Nate Soares · 14m
Further readingfurther reading
The approach of the labsStrategy+
Core Views on AI Safety
Anthropic · ~12m excerpt
An Approach to Technical AGI Safety and Security
Google DeepMind · ~10m excerpt
Our approach to alignment research
OpenAI · 8m
Further readingfurther reading
OptimistsStrategy+
AI is easy to control
Nora Belrose & Quintin Pope · ~15m excerpt
Alignment By Default
John Wentworth · 14m
Unit 6 Current approaches22
Picking a pathCombined U6
The Field of AI Alignment: A Postmortem
John Wentworth · 9m
So You Want To Make Marginal Progress...
John Wentworth · ~8m
AGI safety career advice
Richard Ngo · ~15m excerpt
Further readingfurther reading
Richard Hamming · ~50m e
Agent foundationsCombined U6
The Rocket Alignment Problem
Eliezer Yudkowsky · 19m
Why Agent Foundations? An Overly Abstract Explanation
John Wentworth · 10m
Subsystem Alignment
Abram Demski & Scott Garrabrant · ~15m e
Further readingfurther reading
Technical alignmentTechnical+
AGI Safety Strategies
AI Safety Atlas · ~15m excerpt
What Everyone in Technical Alignment is Doing and Why
Thomas Larsen & Eli Lifland · ~20m excerpt
Zoom In: An Introduction to Circuits
Chris Olah et al. · ~18m
Further readingfurther reading
Ryan Greenblatt & Buck Shlegeris · ~40m
Governance & policyStrategy+
🎬
Rational Animations · 15m
AI Governance: Opportunity and Theory of Impact
Allan Dafoe · ~25m e
Computing Power and the Governance of AI
GovAI · ~12m
Field-building & outreachStrategy+
Building the field of AI safety
80,000 Hours · ~12m e
Strategy & forecastingStrategy+
Forecasting Timelines
AI Safety Atlas · ~12m excerpt
AI 2027
AI Futures Project · ~1h e — the scenario itself
Cooperative & multi-agent safetyStrategy+
What is Cooperative AI?
Cooperative AI Foundation · ~10m e
The Commitment Races problem
Daniel Kokotajlo · 7m
When would AGIs engage in conflict?
Jesse Clifton, Sammy Martin & Anthony DiGiovanni · ~15m
Further readingfurther reading
Jan Kulveit et al. · ~45m e
Course B
Approaches to Solving Alignment
19 Combined 2 Technical+
Unit 1 What are we doing and why?10
Why a theory of agency?Combined U1
The Rocket Alignment Problem
Eliezer Yudkowsky · 19m
Further readingfurther reading
Realism about rationality
Richard Ngo · 5m
Agents, tools, and simulatorsCombined U1
janus · 52m
Will Petillo et al. · ~15m
Will Petillo et al. · ~15m
Further readingfurther reading
Cleo Nardo · 20m
Describing today's systemsCombined U1
Anthropic (Marks, Lindsey & Olah) · 5m
Sympathy for both sides of the egregious misalignment debate
Steven Byrnes · 6m
Further readingfurther reading
Shanahan, McDonell & Reynolds · Nature, ~30m e
Unit 2 How to think about intelligent agents8
Utility and coherenceCombined U2
Coherent decisions imply consistent utilities
Eliezer Yudkowsky · 33m
Coherence of Caches and Agents
John Wentworth · 14m
Further readingfurther reading
Ihor Kendiukhov · 29m
Decision theoryCombined U2
Further readingfurther reading
How agents carve up the worldCombined U2
John Wentworth · 14m
Unit 3 What should agents aim for?10
The target is hard to nameCombined U3
Eliezer Yudkowsky · 7m
The Pointers Problem: Human Values Are A Function Of Humans' Latent Variables
John Wentworth · 14m
Further readingfurther reading
Separation from hyperexistential risk
Eliezer Yudkowsky et al. · and the s-risk literature
Specifying a targetCombined U3
Coherent Extrapolated Volition
Eliezer Yudkowsky · 12m
On Goal-Models
Richard Ngo · 5m
Further readingfurther reading
On the limits of idealized values
Joe Carlsmith · 44m
Learning the targetCombined U3
Cooperative Inverse Reinforcement Learning
Dylan Hadfield-Menell, Stuart Russell, Pieter Abbeel & Anca Dragan · excerpt, ~20m
Value systematization: how values become coherent (and misaligned)
Richard Ngo · ~15m
Value Formation: An Overarching Model
Thane Ruthenis · 43m
Unit 4 Keeping agents on their goals9
Oversight at scaleTechnical+
🎬
How to Align AI: Put It in a Sandwich
Rational Animations · 5m
Geoffrey Irving, Paul Christiano & Dario Amodei · excerpt, ~30m
John Wentworth · 4m
Paul Christiano · 18m
Keeping goals through self-modification (tiling)Combined U4
Tiling Agents for Self-Modifying AI
Eliezer Yudkowsky & Marcello Herreshoff · excerpt, ~20m
Further readingfurther reading
Vingean Reflection: Reliable Reasoning for Self-Improving Agents
MIRI · ~20m e
Abram Demski & Scott Garrabrant · ~15m e
Unit 5 Managing deviations11
CorrigibilityCombined U5
Corrigibility = Tool-ness?
John Wentworth & David Lorell · 12m
A Shutdown Problem Proposal
Elliott Thornley · 8m
Problem of fully updated deference
Eliezer Yudkowsky · 8m
Why Corrigibility is Hard and Important
Raemon · ~21m
Further readingfurther reading
CAST: Corrigibility As Singular Target
Max Harms · excerpt, ~20m e
ControlTechnical+
The case for ensuring that powerful AIs are controlled
Ryan Greenblatt & Buck Shlegeris · ~40m
Limiting optimisationCombined U5
Reframing Impact
Alex Turner · excerpt, ~15m
Quantilizers: A Safer Alternative to Maximizers for Limited Optimization
Jessica Taylor · PDF, ~15m
Further readingfurther reading
Reframing Superintelligence: Comprehensive AI Services
K. Eric Drexler · ~20m e
Thinking Inside the Box: Controlling and Using an Oracle AI
Armstrong, Sandberg & Bostrom · ~20m e
Unit 6 The research frontier27
OrientationCombined U6
Some Summaries of Agent Foundations Work
Alex Flint et al. · 16m
Embedded agencyCombined U6
Abram Demski & Scott Garrabrant · ~12m e
Selection Theorems: A Program for Understanding Agents
John Wentworth · 8m
Scott Garrabrant · 26m
World-models, abstraction & ontologyCombined U6
A Visual Guide to Natural Latents
John Wentworth & David Lorell · 23m
Scott Garrabrant · 30m
Reasoning under uncertainty & bounded rationalityCombined U6
The Learning-Theoretic AI Alignment Research Agenda
Vanessa Kosoy · 40m
A very non-technical explanation of the basics of infra-Bayesianism
David Matolcsi · 12m
Scott Garrabrant · 11m
Scott Garrabrant · 5m
PreDCA
Vanessa Kosoy · ~30m e
Logical & self-referential reasoningCombined U6
Radical Probabilism
Abram Demski · 45m
Multi-agent interactionsCombined U6
Scott Alexander · 8m