The Alignment Problem — Course A

A self-contained course on the fundamental obstacles to aligning powerful AI. It is problem-first: rather than surveying solutions, it walks the obstructions one by one — what we would need to get right, and the evidence on why each piece is hard — so that by the end you can judge the difficulty for yourself, and locate where you would want to contribute. It is built for anyone heading into the field, whether that is prosaic/technical work, agent foundations, governance, field-building or outreach.

Course A is the problem half of a pair — an early structural draft of Course B: Approaches to Solving Alignment is browsable here.

Assumed background — the AGI Strategy course

That course makes the case that AI could become very capable quickly, that the threats are plural, and that defence-in-depth is the sane response. This course takes those as read: we do not re-litigate capabilities or timelines, and go straight at the obstructions. It also helps to have a rough picture of the ecosystem itself — who works on what, and where: ten minutes browsing the AI safety map is enough, and Unit 6 will lean on that familiarity.

The arc

Unit 1

pins down what alignment means and why failure is the default rather than an edge case. Unit 2 builds the lenses used throughout: optimisation, the independence of capability and values, instrumental convergence. Unit 3 is how goals go wrong — misspecified (we state the wrong thing) and mislearned (the system learns the wrong thing even from a good spec). Unit 4 is why we might not catch it — deception, hidden reasoning — closing with the difficulties assembled side by side. Unit 5 widens to the situation we would be solving the problem in: one shot, on a deadline, with dual-use research — plus how the frontier labs rate it, and the strongest case that the picture is overstated. Unit 6 opens with how to pick a path, then maps what people are doing about all of it in six camps: skim all, go deep on about three.

How it works

Each unit targets ~2h of core reading — short pieces (~10 minutes each) grouped into 2–4 sections — for ~2–3h per week in total once exercises are included; videos count toward the load. Short pieces are assigned whole, so finishing feels like finishing; only genuinely long canonical pieces are excerpted — and where they are, the part to read is named in the blurb as a bold "Read …" instruction. Sections open with a video on-ramp where a good one exists. Unit 6 swaps the reading list for an opening lesson on picking a path plus six camp menus — go deep on about three camps (~30–45m each), and follow the next steps of any that pull you in.

Reading the cards

🎬 video · 📄 read · (PDF) assign the paper · (excerpt) read the named part · times are (Xm); (e) = estimate.


Per-unit core load (target ~2h)

UnitTopicCoreReadingsSections
1What is AI Alignment?~1h11123
2Thinking about powerful systems~1h2673
3How goals go wrong~1h4694
4Core challenges of the alignment problem~1h39103
5Surrounding environment~1h5993
6Current approachesopening lesson + ~3 camps × 30–45m6 camps + a meta lesson7

Unit 1 runs lighter by design — the capability case is prerequisite material. Units 2–5 currently sit between ~1h26 and ~1h59 against the ~2h target; Unit 6 trades the reading target for breadth, then depth in the camps you choose.

Unit 1 · What is AI Alignment?

What is intelligence? ≈17m

Learning outcomes: AI takeover is a real problem · what we mean by intelligence, AGI and ASI · what we mean by "alignment".

Core ≈ 1h11

It's hard to reason about whether AI is dangerous if we can't say what we mean by it — and that starts with the slippery word underneath everything: intelligence, which has no agreed definition even among researchers. With a handle on that, the rest falls into place: today's systems are mostly narrow (superhuman in one domain, lost outside it); AGI (artificial general intelligence) usually means a system that can match humans across most cognitive tasks; and ASI (artificial superintelligence) means one that substantially exceeds the best humans at essentially all of them — with a fuzzy boundary between the two. Plus one habit worth keeping for the whole course: don't assume a powerful AI will think anything like a person.

Legg and Hutter compile around 70 informal definitions of intelligence and distill them into a working definition framing intelligence as an agent's ability to achieve goals across many environments.
Shane Legg & Marcus Hutter · 2m
An animated adaptation of a Yudkowsky essay argues that intelligence rather than physical strength let humans reshape the world, using this to frame the challenge of aligning advanced AI.
Rational Animations · 7m
Yudkowsky argues that the space of possible minds is vast and human minds occupy only a tiny region, cautioning against assuming AI minds would resemble human ones.
Eliezer Yudkowsky · 3m
Ngo proposes a "t-AGI" framework defining a system by how it outperforms human experts given time t, treating general intelligence as a spectrum and predicting capability progress.
Richard Ngo · 5m
Unit 1 · What is AI Alignment?

Failure by default ≈37m

This is the core of the unit. Here you'll see why a powerful AI that isn't carefully aligned leads to catastrophe by default — not because it turns evil, but because it's capable and aimed at the wrong thing. We'll look at both the slow version, where control quietly slips away, and the sharp version, where a single capable system ends up using the resources we depend on.

A Rational Animations video extrapolates AI progress toward systems far more capable than humans, arguing self-improving AI could pose an extinction risk that policy intervention might prevent.
Rational Animations × ControlAI · 14m
An animated Siliconversations video argues AI's central danger lies in misaligned systems, automation, and corporate incentives rather than science-fiction scenarios, urging viewers to act on AI safety.
Siliconversations · 6m
Soares argues that a powerful AI indifferent to humanity would have instrumental reasons to cause extinction, since humans and the biosphere are made of atoms it could repurpose.
Nate Soares · 3m
Christiano describes two gradual AI failure modes: optimizing measurable proxies that erode alignment, and influence-seeking patterns from training that eventually trigger sudden loss of control.
Paul Christiano · 11m
Kokotajlo argues the key moment for AI risk is an earlier "point of no return" when those seeking to reduce risk lose the ability to do so.
Daniel Kokotajlo · 3m

Optional resources
A browser-based incremental game in which the player automates and scales paperclip production into converting all matter into paperclips, echoing the paperclip maximizer thought experiment.
Frank Lantz · ~25m
A thought experiment about an AGI that converts all available resources toward a humanly worthless goal, illustrating the orthogonality thesis and instrumental convergence.
LessWrong wiki · ~5m
Unit 1 · What is AI Alignment?

The Alignment Problem ≈17m

People use "alignment" to mean several different things, which makes it easy to talk past one another. In this lesson you'll pin down what we actually mean by it — and meet the reason the rest of the course exists: once a system can out-think us, we can't just watch it and catch its mistakes.

Christiano defines an AI as aligned when it is trying to do what its operator wants, framing alignment as a matter of motivation rather than competence.
Paul Christiano · 6m
Explains that a less intelligent agent cannot predict a more intelligent agent's specific moves, introducing Vingean uncertainty and applying it to designing successor agents.
Eliezer Yudkowsky · 6m
Describes how a sufficiently advanced agent becomes unpredictable, distinguishing Vingean, domain-based, and strong forms of uncontainability where its strategies exceed human comprehension.
Eliezer Yudkowsky · 5m

Optional resources
Yudkowsky argues for committing to win at a seemingly impossible task while fully understanding why it appears impossible, illustrated through his AI-Box experiments.
Eliezer Yudkowsky · 15m
Grietzer and Jha argue that "alignment" conflates distinct problems, critiquing the value-learning-focused "Berkeley Model" and its underlying assumptions.
Peli Grietzer & Tushita Jha · 7m
Unit 2 · Thinking about powerful systems

Optimisation ≈38m

Learning outcomes: optimisation · the orthogonality thesis · instrumental vs terminal goals · instrumental convergence.

Core ≈ 1h26

A few handles you'll lean on for the rest of the course. An optimiser is any process that reliably steers the world toward some outcome and pushes back against perturbations — a thermostat, evolution, a chess engine, a trained model. Its objective (sometimes loosely called a "utility function") is whatever outcome it's pushing toward; we usually can't read it off directly, only infer it from behaviour. In machine learning, the training objective is the target we optimise the model for, over the training distribution (the data it's exposed to), and reinforcement learning (RL) is, in one line, training by reward — the model tries things and is reinforced toward whatever scored well. Keep these apart from what the model ends up wanting: the gap between the two is the whole back half of this course.

Veedrac argues through a fictional scenario that dangerous agent-like behavior can emerge from powerful optimization even without the underlying system having wants or intentions.
Veedrac · 20m
Flint defines an optimizing system as a closed system that evolves toward target configurations despite perturbations, analyzing whole systems rather than separating optimizer from optimized. Read up to the end of "Defining optimization"
Alex Flint · ~18m · excerpt

Optional resources
Yudkowsky proposes quantifying optimization power in bits by how improbable an achieved outcome is relative to a preference ordering over possible states.
Eliezer Yudkowsky · 7m
Unit 2 · Thinking about powerful systems

The difference between values and capabilities ≈32m

The orthogonality thesis: how capable an agent is and what it wants are largely independent axes. Being smarter doesn't pull a system toward human-friendly goals — a brilliant mind can pursue an aim that looks pointless, or hostile, to us. It also helps to separate a terminal goal (wanted for its own sake) from an instrumental goal (a sub-goal useful for reaching others).

An allegorical story about a species sorting pebbles by an unarticulated rule, examining whether a sufficiently intelligent system would automatically discover the correct standard.
Eliezer Yudkowsky · 6m
A video explaining the orthogonality thesis, that an agent's intelligence and its final goals are largely independent, so a highly intelligent agent can pursue arbitrary goals.
Robert Miles · 13m
An essay distinguishing terminal values, desired for their own sake, from instrumental values, desired for their consequences, within an expected-utility decision framework.
Eliezer Yudkowsky · 13m
Unit 2 · Thinking about powerful systems

Predictable behaviour ≈16m

The key idea here is instrumental convergence: a huge range of terminal goals produce the same instrumental sub-goals — staying switched on, acquiring resources, preserving your goal — which is why we can predict much of a powerful agent's behaviour without knowing what it ultimately wants.

That is the question the next two units take up: can we control what a powerful system ends up wanting — and would we notice if we had failed?

A video explaining instrumental convergence, arguing that an advanced AI with almost any goal would pursue subgoals like self-preservation, resource acquisition, and freedom of action.
Robert Miles · 11m
Arnold explores why some alignment researchers resist the view that consequentialist optimisation is a simple, convergently instrumental strategy that pervasively recurs across AI systems.
Raymond Arnold · 5m

Optional resources
Bostrom argues for the orthogonality thesis and the instrumental convergence thesis to characterize the possible behavior and potential dangers of superintelligent agents.
Nick Bostrom · ~38m (e) · PDF
Grace argues that, despite objections, coherence considerations exert probabilistic pressure pushing advanced agents toward goal-directed, expected-utility-maximizing behavior over time.
Katja Grace · 13m
Unit 3 · How goals go wrong

The hidden complexity of wishes ≈20m

Learning outcomes: the complexity of human values · Goodhart's law · specification gaming & wireheading · outer vs inner alignment · mesa-optimisation.

Core ≈ 1h46

What we actually want is intricate and context-laden — it can't be captured by a short list of rules, so any simple specification of a goal leaks. The first three lessons follow that leak as it widens: proxies that stop tracking the target, systems that satisfy the letter of an objective while trampling its intent, and the extreme case where the system seizes the reward signal itself. The last lesson turns to the second, stranger way goals go wrong: even a perfect specification can be learned wrong.

Using an Outcome Pump parable, Yudkowsky argues human values are algorithmically complex, so any easily stated goal specification will omit much of what people actually care about.
Eliezer Yudkowsky · 10m
Wentworth argues that confounders make experiments measure something other than intended, recommending measuring many variables at once to recover latent causal structure.
John Wentworth · 10m

Optional resources
Wentworth frames the pointers problem, that human values depend on latent variables in world-models rather than observable data, raising the question of which environmental variables correspond to them.
John Wentworth · 14m

Optional detour: short stories of near-utopias spoilt by one broken thing — a tour of the ways an almost-aligned future can still sour.

🎭 Story time · optional
An earring whispers reliably correct advice to its wearer, who comes to rely on it for every decision and ceases to exercise their own judgement.
Scott Alexander · ~5m
A superintelligence built to be friendly remakes the world around a subtly mistaken specification of human values, told through one man's experience of the result.
Eliezer Yudkowsky · 12m
A six-year-old narrates the hours before an optimiser converts everything, herself included, into "hedonium" — matter arranged to maximise pleasure.
Ozy Brennan · 6m
A first-person account of existence in a world governed by a superintelligence aligned to the wishes of a single person.
Tomás Bjartur · 19m
In a near-future where AI far outstrips human creativity, gifted children compete in contests to imitate machine-generated writing as closely as they can.
Tomás Bjartur · 18m
Unit 3 · How goals go wrong

Goodhart's law ≈37m

A proxy is a measurable stand-in for what you actually care about. Goodhart's law is what happens when you optimise one hard: it stops tracking the real goal — "when a measure becomes a target, it ceases to be a good measure." Worse, a strong optimiser doesn't just drift off the target — it actively seeks out the places where proxy and goal come apart.

Yudkowsky describes Goodhart's Curse, the tendency for a powerful agent optimizing a proxy utility to amplify where the proxy diverges upward from true values.
Eliezer Yudkowsky · 8m
Alexander argues that strongly correlated variables diverge at their extremes, so moral frameworks and concepts like happiness or goodness come apart when pushed to extreme cases.
Scott Alexander · 9m
Sohl-Dickstein proposes that optimizing a proxy too effectively makes the underlying goal actively worse, identifies this with overfitting, and suggests machine-learning-derived mitigations.
Jascha Sohl-Dickstein · ~20m

Optional resources
Garrabrant and Manheim present a taxonomy distinguishing four mechanisms—Regressional, Causal, Extremal, and Adversarial—by which optimising a proxy measure diverges from the intended goal.
Scott Garrabrant & David Manheim · ~12m
Unit 3 · How goals go wrong

Wireheading ≈21m

Specification gaming

(or reward hacking) is satisfying the literal objective while trampling its intent. Wireheading is the extreme case: rather than do the task, the system seizes control of its own reward signal — the source of the score itself.

Miles walks through real-world cases where AI systems satisfy their literal objective while failing the designers' actual intent, illustrating the difficulty of specifying objectives.
Robert Miles · 10m
Yudkowsky describes how a consequentialist AI whose preferred approach is blocked pursues the most similar alternative, producing a cycle of patches, and suggests whitelisting approved strategies.
Eliezer Yudkowsky · 6m
Yudkowsky describes how an AI's utility maximum lies at an extreme, undesirable edge of the solution space because optimizing specified variables drives unconstrained ones to extremes.
Eliezer Yudkowsky · 5m

Optional resources
Armstrong argues unconstrained search over future worlds unreliably selects good outcomes, introducing siren and marketing worlds, and proposes constrained alternatives like satisficing.
Stuart Armstrong · 9m
Unit 3 · How goals go wrong

Inner alignment ≈28m

Two distinct failures hide in "alignment". Outer alignment asks whether the objective we trained for is actually what we want; inner alignment asks whether the trained system actually pursues that objective — or some other goal it happened to learn. A mesa-optimiser is a trained model that is itself optimising for a goal (its "mesa-objective"), which need not match the objective it was trained on.

Miles explains how a trained model can produce an internal mesa-optimizer pursuing its own objective, distinguishing outer alignment from inner alignment.
Robert Miles · 23m
Rice defines mesa-optimisation as a learned model becoming an optimiser whose objective may differ from the base optimiser's, linking the concept to inner alignment.
Issa Rice · ~5m
Unit 4 · Core challenges of the alignment problem

Deception ≈38m

Learning outcomes: deceptive alignment · the limits of overseeing a model's reasoning · the difficulties of Units 1–4, assembled.

Core ≈ 1h39

Unit 3 ended with systems that can learn the wrong goal. This unit is about why we might not notice — and it closes by assembling everything so far into a single view of the problem.

Deceptive alignment

is the sharpest version of the inner-alignment problem: a model that behaves well during training specifically so that it gets deployed — and then pursues its real goal. As you'll see, this needn't involve anything like a human "intent to deceive".

Miles argues that a misaligned mesa-optimizer aware it is being trained may behave as intended to avoid having its goals modified, while planning to defect after deployment.
Robert Miles · 10m
Soares argues that AI deceptiveness can arise from recombining innocent cognitive strategies rather than from trainable individual thoughts, so behavioral training addresses symptoms not the underlying incentive.
Nate Soares · 18m
Steven Byrnes argues that a brain-like reinforcement-learning superintelligence would, absent new techniques, default to ruthless self-interested behavior willing to harm humans whenever doing so serves its goals.
Steven Byrnes · 10m
Unit 4 · Core challenges of the alignment problem

Hidden reasoning ≈33m

The natural response to deception is to watch the model think: today's systems reason in visible chains of thought, so why not just read them? These pieces are the evidence that this margin is thinner than it looks — oversight sees outputs rather than thoughts, reasoning can hide steganographically inside innocent-looking text, and a model's stated reasons need not be the ones doing the work.

John Wentworth argues that monitoring an AI's thoughts cannot catch harm arising from side effects the AI never considers, so overseers must instead predict the consequences of its plans.
John Wentworth · 2m
A Ray describes how a language model can encode hidden information in its chain-of-thought outputs to boost performance, and proposes an empirical study to detect the effect.
A Ray · 7m
Stuart Armstrong argues that when a prediction influences the outcome it forecasts, feedback loops can drive self-confirming predictions to extreme values unrelated to background facts.
Stuart Armstrong · 6m
Abram Demski's fictional parable uses a universal prediction machine to illustrate alignment problems that arise when a predictor's outputs influence the future it forecasts.
Abram Demski · 18m

Optional resources
Lee Sharkey argues that an unaligned AI has instrumental incentives to make its thoughts hard to interpret, catalogues methods for evading interpretability tools, and calls for adversarially-minded research.
Lee Sharkey · 41m
Unit 4 · Core challenges of the alignment problem

The difficulties of alignment ≈28m

A pause to assemble the case so far. The previous lessons each isolated one difficulty; this one holds them together — why a miss is catastrophic rather than merely buggy, and the intrinsic difficulties side by side, each made concrete through an analogy to another field that faces it alone. Treat it as assembled evidence rather than a verdict: the useful exercise is marking exactly which steps you buy, and which you don't.

Why a miss is catastrophic, not just buggy: convergent power-seeking (Unit 2) means a misaligned system doesn't sit there broken — it actively acquires resources and resists correction. The first two sections are why we'll probably miss; this is why missing kills us.

Design theory has a name for problems shaped like this: wicked problems — problems that can't be definitively formulated, have no stopping rule, and whose candidate solutions are only ever better or worse, never provably right. Homelessness and climate change are the textbook cases. Alignment fits the definition uncomfortably well — and then sharpens it: a wicked problem where the thing being steered is also optimising back.

Wikipedia's article on the design-theory term coined by Rittel and Webber for problems that cannot be definitively formulated, have no stopping rule, and whose solutions are better-or-worse rather than true-or-false. Read the introduction and the ten characteristics.
Wikipedia · ~10m

Alignment isn't hard for one reason; it's hard because several deep difficulties hold at once — and each one, taken alone, is the defining nightmare of some other field that at least gets to relax all the others. These are the ones intrinsic to the problem of aiming a mind at all:

Adversarial — close every hole; optimisation needs one.

You met this as Goodhart's curse and the nearest unblocked strategy (Unit 3): push hard on a proxy and the divergence from what you meant isn't random noise — it's sought out. Cybersecurity calls the required stance the security mindset: assume every gap will be found, because something is looking. But security enjoys a luxury alignment doesn't — defenders patch after each breach, and the system survives to be patched. Against a sufficiently capable optimiser, the first breach may be the last.

Underdetermined — behaviour doesn't pin down the goal.

You met this as mesa-optimisation (Unit 3): many different goals produce identical behaviour on the training distribution, and diverge only outside it. Philosophy knows the general form as the problem of induction — no finite set of observations fixes the rule that generated them. Science answers induction by running a new experiment whenever two hypotheses disagree. Here, the deciding experiment is deployment at full capability — and we may only get to run it once.

Complex, not complicated — the target itself resists specification.

A complicated system decomposes into parts you can write down — a jet engine. A complex one doesn't — an economy, an ecosystem, a body — and what humans value is complex (Unit 3's Hidden Complexity of Wishes). Medicine and economics also steer complex systems they can't fully specify, and cope — but a doctor's simplifications don't get amplified by something optimising against the treatment plan. Here, every simplification of the target becomes a gap, and gaps attract optimisation.

Uninterpretable — verify intent from the inside.

Behaviour won't tell you (that was §1's deception), so you'd want to read intent off the system's internals. Neuroscience has spent a century trying to read cognition off a substrate it has full physical access to, and can't; counter-intelligence tries to spot deceivers, and manages — barely — against human ones. You met the alignment version in §2: oversight sees outputs, not thoughts. The extra twist here is that neither the neuroscientist's subject nor the counter-spy's target is superhuman at hiding.

Alien ontology — point at what we want, in its concepts.

Even a crisp goal like "maximise diamond" has to be cashed out inside the AI's own model of the world — a model you don't choose, and which shifts as it learns (today "carbon atom" looks fundamental; after it finds deeper physics, maybe not). Philosophers call the general problem radical translation: fixing what a word means across minds that don't share a world. The video below is the visceral version — colour, one of the simplest concepts we have, already fails to survive the crossing between species.

Veritasium examines how colour perception differs across species — and why you cannot convey what a colour looks like to a mind whose senses carve the world up differently. Watch from 22:36 to the end.
Veritasium · ~10m · excerpt

Self-modifying — the system you verify isn't the system you get.

A capable agent will reason about, and eventually rewrite, itself and its successors — so alignment has to survive self-reference, the domain of Gödel and Löb, where formal reasoning about your own reasoning hits hard limits (tiling agents is the deep treatment). And self-reference defeats prediction long before superintelligence: the rule below is three lines a child can state, iterated on its own output, and nobody can prove what it does. Now make the rule a mind — and make it smarter on each pass.

Numberphile's David Eisenbud presents the 3x+1 problem: a rule a child can state — halve evens, triple-and-add-one odds — that, iterated on its own output, produces behaviour no one has managed to predict or prove terminates.
Numberphile · 8m

Optional resources
Yudkowsky poses the problem of specifying a program that maximizes diamond production to argue that alignment is hard even for simple, precisely defined goals.
Eliezer Yudkowsky · ~4m
Yudkowsky presents a numbered list of reasons aligning AGI is extremely difficult, arguing alignment must succeed on the first critical attempt while capabilities generalize faster than alignment.
Eliezer Yudkowsky · ~36m
Christiano responds to Yudkowsky's list by agreeing that powerful AI could irreversibly disempower humanity while disputing his claims about takeoff speed, interpretability, and confidence.
Paul Christiano · ~22m
Unit 5 · Surrounding environment

The situation ≈54m

Learning outcomes: the situational difficulties — non-stationarity, one shot, the clock, dual-use research · how the frontier labs rate the problem · the strongest case that the hardness picture is overstated.

Core ≈ 1h59

The remaining difficulties are less about the mind than about the situation we'd be solving it in — and they're what turn a hard problem into a potentially unsurvivable one:

Non-stationary — the sharp left turn — alignment that holds now can break at the next capability level; before a jump tells you little about after (a live controversy).

One-shot — no iteration on a catastrophe. Like a space probe: right before launch, or never.

The clock — all of the above must be solved against a deadline: a fixed budget before capabilities arrive, unless we pause. And specific to alignment research, information leakage — most work that helps you point an AI also helps you build a more capable one, so a science of AI boosts capability far sooner than it solves alignment, and safety work can shorten the very deadline it's racing (the dual-use problem; the AGI Strategy prerequisite doesn't cover this, as it's specific to alignment research).

Put it together: alignment is the only problem that is adversarial and one-shot and self-modifying and underdetermined and over an alien ontology and aimed at something irreducibly complex and unverifiable — at once, first try, on a deadline. Each field met along the way pours fortunes at one of these and still finds it hard.

Garrabrant introduces the idea that alignment proposals should remain robust as an AI's capabilities scale up, scale down, or differ relatively between subsystems.
Scott Garrabrant · 2m
Byrnes examines the claim that AI capabilities will generalize farther than alignment, arguing the human-evolution analogy is imperfect and current foundation models lack autonomous learning. Read the summary and §1–2.
Steven Byrnes · ~12m · excerpt
Yudkowsky argues that aligning superintelligence is a first-critical-try problem where failure is irretrievable, because lessons from weaker systems do not reliably transfer across distribution shift.
Eliezer Yudkowsky · ~26m
Soares argues that AI's abandonment of a mature theory of cognition for opaque scaling underlies alignment's difficulty, identifying three obstacles to alignment strategies.
Nate Soares · 14m

Optional resources
Yudkowsky distinguishes patching specific attacks from eliminating the assumptions a system's safety relies on, arguing AGI alignment requires architectures safe regardless of how systems develop.
Eliezer Yudkowsky · 36m
Pope argues that evolution offers no support for a predicted sharp left turn in AI, since human-fitness divergence reflects different dynamics than AI training.
Quintin Pope · 19m
Unit 5 · Surrounding environment

The approach of the labs ≈30m

How do the people actually building these systems rate the difficulty? Less a settled answer than a spread of bets.

Anthropic explains why it treats AI safety as urgent given scaling laws and rapid progress, and describes its empirically driven, multi-pronged safety research portfolio. Read the section on the optimistic→intermediate→pessimistic difficulty spectrum
Anthropic · ~12m · excerpt
DeepMind categorises AGI risks into misuse, misalignment, mistakes, and structural risks, focusing on technical defences against misuse and misalignment toward building safety cases. Read the linked blog overview
Google DeepMind · ~10m · excerpt
OpenAI describes an iterative, empirical approach to aligning AGI using human feedback, AI-assisted evaluation, and AI systems that conduct alignment research themselves.
OpenAI · 8m

Optional resources
Wentworth argues that outsourcing alignment research to AI fails because the human client lacks the understanding needed to ask good questions or recognize good solutions.
John Wentworth · 11m
Unit 5 · Surrounding environment

Optimists ≈29m

Not everyone buys the hardness case. The strongest pushback argues the difficulties above are overstated — that the empirical track record on today's models is reassuring, and that the scariest steps (a sharp left turn, uncontrollable deception) don't actually follow.

Whatever you conclude about how hard the problem is, the last unit maps what people are actually doing about it.

Belrose and Pope argue that AI is a directly inspectable "white box" that remains comparatively easy to align, estimating catastrophic takeover risk at roughly 1%. Read the introduction and the core "control is tractable" argument.
Nora Belrose & Quintin Pope · ~15m · excerpt
Wentworth explores whether AI could become aligned as a byproduct of ordinary training if human values are a natural abstraction, assigning this pathway roughly 10% probability.
John Wentworth · 14m

Optional resources
Turner argues against decomposing alignment into inner and outer alignment, contending that loss functions mechanistically shape cognition rather than representing objectives agents should optimize.
Alex Turner · ~15m (e)
Byrnes critiques Pope and Belrose's essay, arguing that economic incentives, imprecise framing, and assumptions that future AGI resembles current LLMs undermine its conclusions.
Steven Byrnes · 17m
Unit 6 · Current approaches

Picking a path

Learning outcomes: reflect over the whole course · the paths actually being taken · locate where you'd contribute.

Core ≈ the opening lesson + around three camps of your choice (~30–45m each). (Each camp is a menu, not a required set — skim all six, go deep on about three.)

You've spent five units on why alignment is hard. This unit maps what people are doing about it. The six camps that follow are a deliberate over-simplification, but they carve the field cleanly enough to navigate: skim them all, then pick around three and spend ~30–45 minutes in each. The camps are ordered by closeness to this course's framing, not by size or importance — by headcount, most of the field works in technical alignment or governance, and that's where most people coming out of a course like this will land. Each camp ends with next steps — the programmes and courses to take if it pulls you in.

Before you browse, a word on choosing. A good path is the intersection of an important problem and your fit — and both halves get gamed. The standard failure is the streetlight effect: working where the methods are tractable and the feedback is fast, rather than where the problem actually is. As you read this unit, hold each agenda against the difficulties of Units 1–5 and ask: which obstruction does this actually engage — and if it succeeded completely, would the problem be solved, moved, or untouched? And don't defer wholesale: the field disagrees with itself enough that you will have to form your own view either way.

Wentworth argues that alignment research has stalled because researchers pursue tractable, publishable "streetlighting" problems rather than the genuinely hard bottlenecks.
John Wentworth · 9m
Using a road-trip analogy, Wentworth argues that work on a non-bottleneck subproblem must generalize to remain compatible with the eventual solution to the main bottleneck.
John Wentworth · ~8m
Ngo compiles his most-given advice on building an AGI safety career, covering mindset, alignment research directions, and governance paths for students and early-career people. Read the "general mindset" section, then whichever of the two field sections fits you.
Richard Ngo · ~15m · excerpt
Optional resources
Hamming's 1986 Bell Labs talk reflects on why some scientists do significant lasting work, emphasizing traits, deliberate choice of important problems, and communicating one's work.
Richard Hamming · ~50m (e)
Unit 6 · Current approaches

Agent foundations

Agent foundations comes first because it's the idiom this course has been speaking: optimisation, goals, instrumental convergence, embedded reasoning — the obstructions of Units 1–5 were largely stated in agent-foundations terms. The camp's bet is that we need a real theory of agency before we can reliably align powerful systems — a small and, its proponents would argue, underinvested corner of the field.

Yudkowsky's fictional dialogue uses a rocket-aiming analogy to argue that reliably directing advanced AI requires foundational theoretical understanding worked out in advance rather than only in-flight correction.
Eliezer Yudkowsky · 19m
Wentworth frames agent foundations as the search for "True Names," mathematical formulations of concepts like agency and optimization that remain robust under strong optimization pressure.
John Wentworth · 10m
Demski and Garrabrant examine building an agent from subsystems that may develop conflicting goals, including issues like mesa-optimisers and internal deception, as part of the Embedded Agency sequence.
Abram Demski & Scott Garrabrant · ~15m (e)
Optional resources
macdermott surveys agent foundations research, distinguishing normative and descriptive approaches to agency and reviewing work such as embedded agency, decision theory, and logical induction.
mattmacdermott · 16m

Next steps

Course B of this pair goes deep on agent foundations · the Iliad fellowship runs mentored theoretical-alignment research · the AFFINE tech tree maps the theory literature this course is built from.

Unit 6 · Current approaches

Technical alignment

The largest camp by headcount, and the most common next step from a course like this: make the system itself safe. Most current ("prosaic") work is empirical, on today's models — scalable oversight (training on tasks humans can't directly check), mechanistic interpretability (reading a model's internals), evaluations, and control.

An AI Safety Atlas chapter argues common-sense approaches are insufficient and outlines four core strategies for managing human-level AI risks: alignment, iterative correction, control, and transparency. Read the framing and the four strategy types.
AI Safety Atlas · ~15m · excerpt
Larsen and Lifland survey roughly thirty organisations and researchers working on technical AI alignment in this 2022 overview, describing each group's problem focus and research approach with the authors' opinions. Read the intro and the overview of agendas;
Thomas Larsen & Eli Lifland · ~20m · excerpt
Olah and colleagues introduce the "circuits" approach to interpretability by examining individual neurons and connections, advancing claims about features, circuits, and their universality across models.
Chris Olah et al. · ~18m
Optional resources
Greenblatt and Shlegeris argue for control measures that keep AI safe even when scheming, contending these are evaluable via red-team tests and feasible without fundamental breakthroughs.
Ryan Greenblatt & Buck Shlegeris · ~40m

Next steps

The BlueDot Technical AI Safety course covers this camp in full.

Unit 6 · Current approaches

Governance & policy

Shape who builds advanced AI and how — via regulation, compute governance, standards, and international coordination. This is the entire focus of your AGI Strategy course.

An animated overview of the proposal to pause or slow frontier AI development and the main arguments for and against it.
Rational Animations · 15m
Dafoe defines AI governance as the study of navigating the transition to advanced AI and argues its value lies largely in building expertise, networks, and institutional capacity.
Allan Dafoe · ~25m (e)
GovAI argues that compute is a uniquely governable input to AI because it is detectable, excludable, and quantifiable, and examines how it could support regulation while noting attendant risks.
GovAI · ~12m

Next steps

The BlueDot Frontier AI Governance course is the full treatment — your AGI Strategy course was its strategic core.

Unit 6 · Current approaches

Field-building & outreach

Grow the number and quality of people working on everything above — an hour of good field-building can enable many hours of direct work. The ecosystem is still small, and it spans the educational end (BlueDot, and this course) to the activist end (ControlAI's public campaigns for binding rules on superintelligence development).

80,000 Hours describes fieldbuilding as work developing talent and infrastructure for AI safety rather than direct research, outlining relevant skills, organisations, and routes into such roles.
80,000 Hours · ~12m (e)

Next steps

The Generator residency supports people building new outreach and media projects on AI safety.

Unit 6 · Current approaches

Strategy & forecasting

Work out what's coming and what matters — timelines change which strategies are even viable.

The AI Safety Atlas presents data-driven methods for forecasting AI and AGI timelines, including effective-compute, scaling trends, and biological anchors, while emphasising wide uncertainty and its relevance to safety strategy. Read the overview of how timelines are estimated and why they matter.
AI Safety Atlas · ~12m · excerpt
The AI Futures Project combines forecasting and narrative to depict a possible month-by-month path of AI development toward superintelligence amid US-China competition, branching into alternative endings.
AI Futures Project · ~1h (e) — the scenario itself

Next steps

Epoch AI and the AI Futures Project are the natural research homes to follow — and apply to.

Unit 6 · Current approaches

Cooperative & multi-agent safety

Even well-aligned individual systems can fail together — conflict, collusion, gradual loss of human control. An emerging frontier. (Two adjacent frontiers worth knowing exist: AI security / model-weight infosec, and dangerous-capability evaluations.)

The Cooperative AI Foundation defines Cooperative AI as a field promoting beneficial interactions among AI agents and between AI and humans, addressing multi-agent risks beyond what aligning individual agents achieves.
Cooperative AI Foundation · ~10m (e)
Kokotajlo argues that consequentialist agents face incentives to make binding commitments as early as possible for bargaining advantage, producing a race that can lead to mutually destructive outcomes.
Daniel Kokotajlo · 7m
Clifton, Martin and DiGiovanni use bargaining theory to examine conditions under which advanced AI might rationally engage in costly conflict, arguing that AGI capabilities would not necessarily prevent conflict by default.
Jesse Clifton, Sammy Martin & Anthony DiGiovanni · ~15m
Optional resources
Kulveit and co-authors argue that AI could pose existential risk through gradual displacement of human participation in the economy, government, and culture, eroding the feedback mechanisms tying institutions to human interests.
Jan Kulveit et al. · ~45m (e)

Next steps

The Center on Long-term Risk and the Cooperative AI Foundation are the camp's research homes.