AGI Strategy — Racing to a Better Future · BlueDot (editable mockup)

A local, editable mockup of BlueDot's AGI Strategy course, scraped from bluedot.org for restructure experiments — this is not the canonical BlueDot course, and reading blurbs/intros are BlueDot's own text. Use it to try folding in the offloaded A/B material.

Reading the cards

🎬 video · 📄 read · times are (Xm); tiers: Core readings shown inline, others under "Optional resources".


Per-unit core load

UnitTopicCoreReadingsSections
1Racing to a Better Future1h4054
2Drivers of AI Progress2h2593
3The Alignment Problem2h17114
4Pathways to Harm4h45196
5Defence in Depth1h4065
6Start Contributing2h30107
Unit 1 · Racing to a Better Future

Imagining a better future ≈55m

Core ≈ 1h40

🎬 Watch the embedded video

You’re the product of 8,000 generations of humans. You were born into the civilization they built, with all its magic, beauty and flaws. Their decisions shape every aspect of our world.

We’re only the latest generation of humans, but we’re living through one of the most significant technological transformations in the history of the universe.

Our decisions today will have an immense impact on our own future, and the future of all our descendants.

We may still be early in humanity’s story. We have an amazing opportunity to steer this story towards wonder, greatness and kindness. But that's not guaranteed.

In this unit, you’ll visualise a future to aspire towards, you’ll explore how AI is shaping that future, and you’ll analyse the dilemmas and coordination failures we must overcome to steer towards a better future.

Throughout the rest of the course, you will:

analyse the technical trends of AI progress and what this implies for future AI capabilities,

develop a step-by-step story (”kill chain”) for how those capabilities could cause harm to humanity,

apply a “defence in depth” framework to analyse what needs to be built to defend against your kill chain and steer AI towards safe and beneficial outcomes,

generate your own action plan for how you'll start contributing to making AI go well.

The Institute for Progress lays out how the US Government could shape the development of AI towards human flourishing by accelerating beneficial AI applications and defences against societal harms.
Tim Fist, Tao Burga, and Tim Hwang · 30m
Read chapter 1: The Return of Utopia Historian Rutger Bregman demonstrates that humanity's greatest achievements—from democracy to eradicating smallpox—began as dismissed utopias, and argues we need to reclaim this visionary thinking to build the future we want.
Rutger Bregman · 25m
Unit 1 · Racing to a Better Future

What future do you want?

Exercise — What does a better future for yourself look like?.

In so many ways, human lives have been transformed in the past few centuries. In many ways, perhaps most, it’s been for the better. However, there’s a huge amount of work left to be done to steer our future towards wiser, kinder, happier outcomes.

How old are you? Add 20 years to that number. Imagine a week in your life when you’re that years old, where you’re living a good, happy life. Consider going for a short walk to reflect on this. Then write down your thoughts below.

_We appreciate that this is quite a personal question! In general, your exercise responses are accessible to the BlueDot team and your group facilitator. If you would rather note down your thoughts in a private doc then you're very welcome._

Exercise — What does a better future look like for society?.

Assuming current population trends, over the next 20 years we will welcome 2.7 billion new humans to our global community, and 1.5 billion will pass away.

If you felt pride in welcoming these new humans into our world, what would the world look like?

Consider factors like politics, the economy, the social order and community, technological progress, and the environment.

Unit 1 · Racing to a Better Future

Steering the race to AGI ≈45m

You've just imagined futures worth protecting. But those futures aren't guaranteed - they depend in part on the decisions being made right now by the people racing to build AGI, who often have incentives that push away from thoughtful development.

Here, you’ll learn more about the actors shaping the future of AI, including their goals and constraints.

This essay describes AI safety policies that rely on centralised control (surveillance, fewer AI projects, licensing regimes) as "stasist" approaches that sacrifice innovation for stability. Toner argues we need "dynamist" solutions to the risks from AI that allow for decentralised experimentation, creativity and risk-taking.
Helen Toner · 15m
Archived from It’s practically impossible to run a big AI company ethically (Vox). Even "safety-first" AI companies like Anthropic face market pressure that can override ethical commitments. This article demonstrates the constraints facing AI companies, and why voluntary corporate governance alone can't solve coordination problems. Note that we don't intend to critique Anthropic specifically: all AI companies face similar market and ideological pressures, and Anthropic did endorse a recent AI transparency bill (SB 53).
Sigal Samuel · 15m
This RAND article describes some of the international dynamics driving the race to AGI between the US and China, and analyses whether nuclear deterrence logic applies to this race.
Rehman, Mueller, and Mazarr · 15m
Unit 1 · Racing to a Better Future

The characters

Exercise — The actors shaping the future of AI.

Choose one character from this list, and read their character card.

If you're completing this course with a group, post in Slack which character you selected.

Write down 2-3 paragraphs below describing your character's main motivations, capabilities, and external pressures/constraints.

Note that these cards are illustrative - they do NOT capture the true complexity of these characters.


Optional resources
Anthropic's CEO, Dario Amodei, spends most of his time worrying about AI catastrophes. In this essay, he makes the opposite case: if we avoid the biggest risks, powerful AI could eliminate most disease, drastically reduce poverty, and strengthen democracy by compressing a century of progress into a single decade.
Dario Amodei · 75m
This blog post offers a vivid, optimistic vision of rapid AI progress from the CEO of OpenAI. Altman suggests that the accelerating technological change will feel "impressive but manageable," and that there are serious challenges to confront.
Sam Altman · 10m
What might sustainable human progress look like, beyond pure technological acceleration? This essay provides an alternative vision, based on communities living in greater harmony with each other and with nature, alongside advanced technologies.
Joshua Krook · 10m
OpenAI's CEO argues AI will create unprecedented prosperity by giving everyone personalised AI assistants for education, healthcare, and problem-solving—from virtual tutors to breakthroughs in climate and physics—though he warns this depends on building enough computing power to keep AI accessible rather than scarce.
Sam Altman · 5m
Unit 2 · Drivers of AI Progress

Technical trends driving AI progress ≈30m

Core ≈ 2h25

What should we do if superhuman intelligence becomes "too cheap to meter"?

In this unit, we examine why AI capabilities keep getting better: cheaper and more plentiful compute, and better algorithms. We then analyse what that implies for the rate of future AI capabilities, and how quickly those capabilities will spread to many actors all over the world.

Definitions:

FLOP: a single basic mathematical operation (like addition or multiplication) that a computer performs. The total number of FLOP is a measure of computational work done during AI training.

Compute efficiency is a measure of how “good” an AI model you get with a given financial investment.

E.g. if OpenAI spends $500M on compute to train their next AI model, how "good" will it be?

Hardware price-performance measures the amount of computational resources (FLOP) available for a given financial investment.

E.g. how many FLOP would OpenAI get with a $500M investment into NVIDIA H200 AI chips running for 1 month?

This is improving at ~1.4x/year (i.e. you get 1.4x more FLOP per $ each year) due to Moore's Law.

Algorithmic efficiency measures how “good” of an AI model you get with a given number of FLOP.

E.g. if OpenAI run their NVIDIA H200 chips for 1 month training an AI model (~10^26 FLOP), how “good” will that AI model be?

This is improving at ~3x/year (i.e. you need 3x less FLOP to achieve the same capability each year).

Compute efficiency (Capability/$) = Hardware price-performance (FLOP/$) × Algorithmic efficiency (Capability/FLOP)

This is improving at ~4x/year (i.e. you need 4x less $ to achieve the same capability each year)

Read the executive summary. "Machine learning systems use computing power to execute algorithms that learn from data." This piece introduces the "AI Triad", the three main ingredients for training AI systems: 1. Algorithms 2. Data 3. Computing power
Ben Buchanan · 5m
This post explains the "scaling laws" that drive rapid AI progress: when you make AI models bigger and train them with more computing power, they get smarter at most tasks. The piece also introduces a second scaling law, where AI performance improves by spending more time "thinking" before responding.
Ethan Mollick · 15m
Use these graphs to analyse how much compute is being used to train frontier AI models, and how quickly those capabilities diffuse due to improvements in hardware and algorithms.
Dewi Erwan · 10m
Unit 2 · Drivers of AI Progress

Intelligence Explosion ≈1h12

> Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an 'intelligence explosion.'

— Irving J Good, 1965

An animated adaptation of a Yudkowsky essay arguing that it was intelligence, not physical strength, that let humans reshape the world — a vivid sense of how much more capable a more intelligent agent can be.
Rational Animations · 7m
This blog post traces AI's rapid leap from high school to PhD-level intelligence in just two years, examines whether physical bottlenecks like computing power can slow this acceleration, and argues that recent efficiency breakthroughs suggest we're approaching an intelligence explosion.
Tomas Pueyo · 30m
Tim Urban uses historical analogies to show why AI progress might accelerate much faster than we expect, and how AI systems could rapidly self-improve from human-level to superintelligent capabilities.
Tim Urban · 25m
Helen Toner, former OpenAI board member, reveals how the AI timeline debate has compressed: even conservative experts who once dismissed advanced AI concerns now predict human-level systems within decades. Rapid AI progress has shifted from a fringe prediction to mainstream expert consensus.
Helen Toner · 10m
Unit 2 · Drivers of AI Progress

Will AI progress accelerate even faster? ≈50m

Helen Toner's piece asks 3 fundamental questions about the trajectory of AI development. These are 1. How far can we get with the current AI paradigm? 2. How much AI can improve itself? 3. Will future AIs will be tools we use, or agents that can act without us? The answers to these questions will determine whether, and how soon, transformative AI will arrive.
Helen Toner · 10m
This piece presents the main arguments for and against the arrival of transformative AI within the next decade.
Sarah Hastings-Woodhouse · 15m
The authors of the two most popular recent AI forecasts (AI 2027 and AI as Normal Technology) discuss where they agree. Surprisingly, there is much common ground. Both sides agree that AI will be at least as transformative as the internet, and both agree that the problem of getting AIs to always act in ways we want (the alignment problem) is unsolved.
Kapoor et. al. · 25m

Exercise — Why are AI capabilities advancing at such a rapid rate?.

Test your understanding of the content. Don’t use jargon or fancy words — provide your explanation in simple english.

Start by defining these terms from memory:

FLOP

Compute efficiency

Hardware price-performance

Algorithmic efficiency

Before continuing, go back to the materials and compare your response.

Exercise — What are the implications of rapid AI progress?.

First, what are the trends driving rapid AI progress?

Second, what are the economic, social, technological and geopolitical implications of rapid AI progress? What impact might this technology have on the world?


Optional resources
OpenAI's leadership outline how humanity might govern superintelligence, proposing international oversight with inspection powers similar to nuclear regulation. They argue the AI systems arriving this decade will be "more powerful than any technology yet created" and their control cannot be left to individual companies alone.
Sam Altman, Greg Brockman, Ilya Sutskever · 5m
Arvind Narayanan and Sayash Kapoor
Unit 3 · The Alignment Problem

Thinking about intelligent systems ≈31m

Core ≈ 2h17

In the past couple of units, you have thought about what a good future might look like and seen that AI may play a large role in it.

In this unit, we take a look at some of the challenges faced when attempting to create an AI which will help rather than hinder.

Before diving into the readings, you may find the following definitions helpful:

Definitions:

Terminal goals: an end-state desired for its own sake.

E.g. you have no ulterior motive behind wanting to be happy: you just want to be happy.

Instrumental goals: an end-state pursued as a means to a further end.

E.g. you work in order to acquire money, but you probably do not value money as an end in itself: its value comes from enabling you to buy a house, buy food, or pay for holidays. More broadly, we would say that acquiring money is instrumentally useful to the terminal goal of being happy.

Orthogonality thesis: the statement that more or less any terminal goal can in principle be combined with any level of intelligence.

E.g. it is both possible to be extremely capable at influencing the world, and to use this capability entirely in the service of making paperclips (see paperclip maximiser).

Instrumental convergence: intelligent agents with different terminal goals still seem to converge on similar instrumental goals. Instrumental Convergence states that this holds for almost any terminal goal, however strange.

E.g. making money is broadly useful within today's society, and people with a wide range of different interests pursue it. Other examples of goals which are broadly instrumentally useful: power, influence, resource acquisition, freedom (i.e. lack of restrictions on the actions one can take).

Paperclip maximiser: an illustrative theoretical superintelligent AI with the sole goal of making as many paperclips as possible. Through instrumental convergence, in the standard telling, the agent cooperates with its supervisors, acquires money and political influence, and bides its time until it can escape its restrictions. At this point, as AI Safety researcher Eliezer Yudkowsky put it, "the AI does not hate you, nor does it love you, but you are made out of atoms which it can use for something else". The scenario usually ends with the observable universe being turned into paperclips.

It is also illustrated in the popular 2017 clicker game Universal Paperclips, which puts the player in the position of the paperclip maximiser.

Utility function: a function which takes in a state of the world and outputs a number reflecting how desirable this world is to a particular agent. This framework is sufficiently general that in principle the preferences of any agent, so long as they are coherent, should be expressible as a utility function.

E.g. the paperclip maximiser could have a utility function U = #Paperclips in the universe.

Reinforcement learning (RL): a way of training AI systems by reward rather than instruction: the system tries actions, receives a numerical reward for the results, and is adjusted to get more reward over time.

The Alignment Problem: the problem of getting a powerful AI system to try to do the things its operator wants it to do.

An allegorical story about a species that sorts pebbles into "correct" heaps by a rule it never articulates, asking whether a smarter mind would discover the rule is right — or just sort faster.
Eliezer Yudkowsky · 7m
The general argument for the orthogonality thesis: an agent's intelligence and its final goals are independent axes, so a highly capable system can in principle pursue almost any goal.
Robert Miles · 13m
Presents the argument that an advanced AI with almost any final goal would pursue sub-goals like self-preservation, resource acquisition, and avoiding being switched off.
Robert Miles · 11m
Optional resources
Armstrong's paper-length defence of the orthogonality thesis against the main counterarguments.
Stuart Armstrong · 35m
The 2008 paper that introduced the instrumental-convergence argument in its original form: Omohundro argues that sufficiently capable goal-directed systems of almost any design will exhibit the same "basic drives" — self-improvement, rationality, goal- and reward-integrity, self-protection, resource acquisition.
Steve Omohundro · ~37m · PDF
Unit 3 · The Alignment Problem

Troubles with specification ≈35m

Given that AIs can have any goal, and giving them the wrong one can be problematic, it's worthwhile to think carefully about how we might specify what we want.

This section explores some of the difficulties that can arise when attempting to do this.

Using a genie-like "Outcome Pump", Yudkowsky argues that any easily stated wish — "get my mother out of the burning building" — omits most of what you actually care about, with catastrophic literal readings. Prefer video? Watch the animated version (11m).
Eliezer Yudkowsky · 14m
Documented cases where AI systems satisfied their literal objective while failing the designers' intent; the authors argue the problem gets harder as systems get more capable. Prefer video? Watch Rational Animations' version (10m).
DeepMind · 8m
Measures that agree over the normal range come apart at the extremes, told through examples from morality and everyday life.
Scott Alexander · 13m

Exercise — The genie test (~10m).

Write down, in two or three sentences, an instruction you'd give a superintelligent assistant managing your household. Then reread it as a literal-minded optimiser with no common sense and superhuman competence: what's the cheapest way to satisfy your exact words, and what does it destroy that you cared about? Revise once and repeat. Notice how the patches multiply.

Optional resources
Unit 3 · The Alignment Problem

Can't we just…? ≈18m

So, specifying what goal we want our AI to have is hard, but can't we just…?

Argues that a sufficiently advanced agent becomes unpredictable in a specific sense, distinguishing Vingean, domain-based, and strong forms of uncontainability where its strategies exceed human comprehension.
Eliezer Yudkowsky · 3m
Yudkowsky argues that a goal-directed AI whose preferred approach is blocked will pursue the most similar alternative, producing a cycle of patches, and considers whitelisting approved strategies as a response.
Eliezer Yudkowsky · 12m
Short answers to the most common "can't we just" proposals, with links onward to fuller treatments of each. Browse the linked answers and read the two or three closest to your own "can't we just…" (reading all of the linked answers takes ~20m).
aisafety.info · ~3m + ~10m browse
Optional resources
Works through the argument that naive ways of giving an AGI an off-switch fail: depending on how its reward is set up, the agent is incentivised either to stop you pressing the button or to press it itself.
Robert Miles (Computerphile) · 20m
Arbital page arguing that building an agent that genuinely lets you correct or switch it off is unnatural for a goal-directed system — the theory behind the off-switch problem.
Eliezer Yudkowsky · ~5m
Unit 3 · The Alignment Problem

Alignment with deep learning systems ≈33m

Many of the arguments that you have seen so far in this unit date from the 2000s and 2010s, long before ChatGPT existed. Here, we look at why these ideas are still relevant to the trajectory we are on today.

Ajeya Cotra argues, for a general audience, that deep learning produces opaque models whose motivations we can't inspect — illustrated by three model types (Saints, Sycophants and Schemers) that would behave identically in training but diverge after deployment. (Where she says "PASTA", read: an AI system able to automate scientific and technological advancement.)
Ajeya Cotra · ~25m
Byrnes argues that pessimists and LLM-focused optimists about AI misalignment are often discussing different systems, and that the niceness of current LLMs is weak evidence about future superintelligence.
Steven Byrnes · 8m
Optional resources
Soares' direct answer to the blunt question: he argues that even a powerful AI merely indifferent to humanity would have instrumental reasons to remove us, because we sit on resources it can use and we might turn it off.
Nate Soares · 3m
Unit 4 · Pathways to Harm

Map the threat landscape ≈55m

Core ≈ 4h45

The future could be amazing. You described it yourself in the first unit. That future might not happen by default. We have a tremendous opportunity and responsibility to make that future a reality for ourselves, everyone we know and love, and the rest of humanity.

If we don’t know what threatens that future, we’re flying blind.

Here, we’ll offer a sketch of the territory. This includes the things we’re protecting, who or what might attack them, and how they might try.

This is not intended to be a perfect overview: treat it like a starting point for your thinking.

The Future of Life Institute show how existing AI can already help to design bioweapons, amplify cyberattacks, and deceive human. Each threat is backed by existing evidence. This gives you concrete examples to draw from when mapping your own threat scenarios.
Ben Eisenpress · 5m
In this talk, Richard Ngo argues that "misalignment" and "misuse" are two sides of the same coin, and governance and technical interventions against examples of misalignment and misuse are frequently the same.
Richard Ngo · 5m
Holden Karnofsky (member of technical staff at Anthropic, founder of GiveWell and the philanthropic foundation Open Philanthropy) describes how human-level AI could spawn hundreds of millions of copies of itself that coordinate to physically defeat humanity. They might accumulate resources, recruit allies, and develop weapons to overpower civilisation. This provides a concrete pathway showing how AI doesn't need to be superintelligent to threaten humanity; it just needs to outnumber us while working as a unified force against us.
Holden Karnofsky · 25m
Security experts document how AI chatbots can now guide non-experts through creating bioweapons from readily available ingredients, eliminating the technical complexity that has historically prevented terrorists and apocalyptic groups from mounting successful biological attacks.
Kyle Hiebert · 5m
Read the summary and the intro. In this blog post, researchers at the Forethought Foundation describe how advanced AI could enable a single person to overthrow governments without any human supporters, through three mechanisms: military AI systems programmed with unwavering loyalty to one leader, secretly loyal AI that passes security tests but executes coups when deployed, and monopolistic access to superhuman capabilities in weapons development and cyber warfare.
Tom Davidson, Lukas Finnveden, Rose Hadshar · 15m
Unit 4 · Pathways to Harm

Prioritising threat pathways ≈5m

In this section, you will choose ONE threat pathway "option" to investigate further - note that your choice will carry over to Unit 5, where you will think about defenses to your chosen threat pathway.

Here's a list of actors we might be concerned about, in terms of them having the capability and/or the motivation to use AI to cause harm to humanity.

"Misaligned AI", i.e. AI systems that act against the interests of humanity.

Powerful human actors, e.g. corporate CEOs, military and political leaders.

Malevolent nation states, e.g. North Korea.

Terrorist groups and doomsday cults.

How might they cause catastrophic harm?

They could use information warfare and conventional military force to overthrow existing power structures like democratic governments.

They could do cyberattacks on critical infrastructure, including the water, energy and food systems.

They could design, build and release viruses into the population that are worse than SARS-CoV-2, leading to a global pandemic.

Or the default incentives for all actors lead to bad outcomes, even absent any malicious intent.

Bruce Schneier · 5m

Exercise — Zooming into one threat.

Based on your current understanding, which threat pathway do you think is of most concern?

Consider the combination of likelihood of the threat, and the impact it would have if it happened.

We don't expect you to have detailed, well-thought out thinking on this at this stage - consider this the beginning of your thinking.

Exercise — Write your own threat scenario.

For your chosen threat pathway, write a "threat scenario" sentence, using this template:

> The [ACTOR] with [CAPABILITY] and [MOTIVATION] attacks [ASSET] by [ATTACK PATHWAY] in order to [OBJECTIVE].

Your exercise response to prioritise which of the next 3 sub-units to explore in more depth.

Unit 4 · Pathways to Harm

Option 1: Power concentration ≈50m

A small group gains control of frontier AI systems and uses them to manipulate elections through micro-targeted disinformation, fake evidence, and synthetic media so convincing that truth becomes unknowable. They deploy AI-powered surveillance that tracks every digital interaction, predicts dissent before it emerges, and automatically suppresses opposition through economic and social levers.

Within months, autonomous weapons patrol the streets, facial recognition determines access to resources, and AI advisors whisper into the ears of every official - all controlled by those who seized the models first. Democratic institutions become theatrical performances while real power flows through algorithmic systems accountable to no one.

Whether it's a government using AI to become permanently unremovable, a corporation achieving regulatory capture through superintelligent lobbying, or an AI system itself accumulating power while maintaining a facade of servitude, the result is the same: the permanent concentration of control in hands that will never voluntarily let go.

Read section 4 on "Concrete paths to an AI-enabled coup".
Tom Davidson, Lukas Finnveden, Rose Hadshar · 15m

Exercise — Create a step-by-step breakdown of the threat.

A kill chain takes a ""threat scenario"" and breaks it into stages of execution. It traces how an attacker would actually proceed step by step, e.g. via reconnaissance, delivery, exploitation, persistence, action on objectives.

Your scenario describes what might happen, whereas a kill chain describes how it would unfold in practice. This enables us to spot choke points where defenders can intervene.

Based on your scenario, use this template to complete this exercise.

Unit 4 · Pathways to Harm

Option 2: Gradual disempowerment ≈1h05

As AI systems become cheaper and more capable than humans at nearly every task, companies face overwhelming pressure to automate. First it's customer service and data entry, then management and creative work, eventually even governance and resource allocation.

This could lead to humans losing economic leverage as their labour becomes worthless. Our ability to influence culture might degrade as AI generates all the content we consume. Our political power could dissipate as governments pivot to AI for their tax revenue, administration, and security, instead of human income tax, human civil servants, and human military personnel.

We might wake up one day to find ourselves comfortable but irrelevant - living in a world we no longer understand or control, where the economy optimises for AI activities rather than human needs, and our preferences matter as little as those of retired plough horses.

Read sections 3 and 4.
Luke Drago and Rudolf Laine · 35m
Aniket Chakravorty, Dewi Erwan · 15m
Samuel Hammond · 15m

Exercise — Create a step-by-step breakdown of the threat.

We now want you to develop your threat scenario into a step-by-step account of how the threat might manifest.

Gradual disempowerment is unlike the other threat models in this unit, because it does not involve a specific “bad actor” attempting to bring about a catastrophic outcome. Instead, humanity is disempowered because of the sum of seemingly benign choices. As a result, we want you to create an “Erosion Pathway/engage in failure mode mapping”. This is a step-by-step guide to how gradual disempowerment would unfold in practice. Based on your scenario, use this template to complete this exercise.

Unit 4 · Pathways to Harm

Option 3: Catastrophic pandemics ≈45m

Advanced AI could dramatically lower the barriers to creating pandemic pathogens. Today, designing a novel virus requires rare expertise, expensive equipment, and years of work. In a few years, AI could walk someone through synthesising a pathogen more transmissible than measles and deadlier than rabies.

Unlike natural pandemics that evolve toward transmission over lethality, an engineered pathogen could be optimised for maximum harm - combining long asymptomatic periods with high mortality, resistance to existing treatments, or even targeting specific genetic markers.

The same AI capabilities that accelerate beneficial biomedical research - protein folding prediction, drug discovery, genetic engineering - could be misused by state bioweapons programmes, terrorist groups, or disturbed individuals to create humanity's final pandemic.

Skim pages 88 to 103.
Anthropic · 15m

Exercise — Create a step-by-step breakdown of the threat.

A kill chain takes a ""threat scenario"" and breaks it into stages of execution. It traces how an attacker would actually proceed step by step, e.g. via reconnaissance, delivery, exploitation, persistence, action on objectives.

Your scenario describes what might happen, whereas a kill chain describes how it would unfold in practice. This enables us to spot choke points where defenders can intervene.

Based on your scenario, use this template to complete this exercise.

Unit 4 · Pathways to Harm

Option 4: Critical infrastructure collapse ≈1h05

Modern infrastructure runs on outdated industrial control systems and unpatched software that were never designed for internet connectivity. AI could systematically map these vulnerabilities across utilities, hospitals, and supply chains, then coordinate attacks that would take human hackers months to plan and execute.

We've already seen previews: the Colonial Pipeline shutdown disrupted fuel supplies across the US East Coast, NotPetya malware accidentally escaped its target and caused $10 billion in global damage, and ransomware regularly forces hospitals to divert ambulances and cancel surgeries. But these were relatively simple attacks with limited coordination.

AI-enabled attackers could compromise the industrial control systems that run power plants, the SCADA networks that manage water treatment, and the logistics software that routes food and medicine - not enough to collapse civilisation, but enough to cause sustained disruptions that compound over weeks. When the power grid has rolling blackouts, water pressure drops, hospitals run on backup generators, and just-in-time supply chains fail, even a wealthy nation starts to fray at the edges.

This resource documents real attacks on power grids and water systems to show the core vulnerability of critical infrastructure: interconnected control systems (computers managing physical equipment) mean a single breach can cascade across sectors, with estimates of over $240 billion in insured losses from one coordinated attack.
Allianz Insurance · 10m
AI is eliminating the expertise and time barriers that have historically made cyberattacks on critical infrastructure difficult. This resource explains how attacks that once took nation-states years to plan can now be automated and executed in minutes, potentially outpacing human defenders
Li-Lian Ang · 5m
This post walks through exactly how hackers might use AI to take down the electricity grids serving millions of people, showing that whilst such attacks remain extremely difficult today, AI tools are already helping hackers work faster and more autonomously.
Byron Tomes · 30m
An interactive demo on how AI can generate personalised phishing emails for any target in minutes, making the attack method that starts 91% of cyber breaches, including those on critical infrastructure like hospitals and power grids, cheap enough to scale to everyone.
CivAI · 10m

Exercise — Create a step-by-step breakdown of the threat.

A kill chain takes a "threat scenario" and breaks it into stages of execution. It traces how an attacker would actually proceed step by step, e.g. via reconnaissance, delivery, exploitation, persistence, action on objectives.

Your scenario describes _what might happen_, whereas a kill chain describes _how it would unfold in practice._ This enables us to spot choke points where defenders can intervene.

Based on your scenario, use this template to complete this exercise.


Optional resources
CrowdStrike, an American cybersecurity company, explains how attackers use AI to automate reconnaissance, craft personalised phishing at scale, and adapt in real-time to evade detection - capabilities that make infrastructure attacks faster and harder to stop than traditional cyberattacks.
Crowdstrike · 10m
Dario Amodei · 100m
Kulveit et al. · 60m
Read the intro and Part 1.
Paul Christiano · 10m
Produced by two course alumni: Stuart Wilmot & Callum Kindred
Future Atlas · 20m
Unit 5 · Defence in Depth

What might success look like?

Core ≈ 1h40

One way to break down existing AGI Strategies is into the following three broad buckets.

This is a gross over-simplification, but we believe it captures the essence of the main “camps” for how to make AI go well.

1: Government control over AGI

This strategy argues in favour of centralising control over compute and frontier AI model weights into a small number of Government-backed “AGI Projects”. These projects would either coordinate with each other to ensure that one doesn’t race ahead to build a dangerous AI system, and/or they’d keep each other in check via coercion, surveillance and strategic deterrence.

Prominent proposals of this strategy are “The Project” by Leopold Aschenbrenner, “CERN for AI”, “Chips for Peace” by Cullen O’Keefe, "Technical Requirements for Halting Dangerous AI Activities" by Barnett et al., and “Superintelligence Strategy” by Hendrycks et al.

In this strategy, global development of frontier AI systems is controlled by these small number of actors, who only build more capable AI systems when they’re confident that doing so would be safe and they can maintain control over the AI system. The project would need to be very secure, both from within (so the AI model itself does not escape, or get helped to escape via internal accomplices), and from external actors trying to steal the AI model weights.

Information about algorithmic innovations might need to be made “born secret”, i.e. top secret/confidential from the point at which they’re created, to prevent actors with smaller amounts of compute from training frontier AI systems in the future.

Proponents of this strategy believe that current geopolitical and commercial race dynamics will push towards the development of AI systems that could cause human extinction. They don’t believe it’s possible to build an “aligned superintelligence”, at least in the short-term, and they don’t believe it’s possible to build defences against a “misaligned superintelligence”. They believe frontier AI development must be stopped, paused or tightly controlled by the world’s governments.

Concerns about this strategy include that it could lead to an extreme concentration of power among the groups controlling the project(s), that it’s intractable and undesirable for Governments to take control over global AI compute supplies, and that it would fail to prevent other actors from building dangerous AI systems anyway due to algorithmic progress.

2: Hand over control to aligned superintelligence

This strategy argues that building superintelligence is inevitable, and that only a superintelligent AI could steer us towards a utopian future and protect humanity from harm. Proponents believe that “AI alignment” is possible, i.e. building AI systems that reliably take actions that follow the interests of the AI company that’s trained it. (Note that AI alignment means different things to different people, spanning from “does what this specific user wants” to “acts in line with human values” to “doesn’t try to kill everyone”).

Actors pursuing this strategy want the “good guys” to “win” the race to superintelligence. Some proponents of this strategy argue that the aligned superintelligence should take advantage of its superior intellectual capabilities to gain a “decisive strategic advantage” over all other AI projects, in order to achieve some form of world domination, and to prevent anyone else from building a “misaligned” superintelligence.

This strategy isn’t made explicit by many actors (”put the AI’s in control of the universe” doesn’t have wide appeal!), but it is the implicit strategy of many actors in the AI ecosystem. To hear this for yourself, go to San Francisco house parties with AI company employees.

You can hear a hint of this here in an interview with Anthropic co-founder Tom Brown. > There’s going to be a handoff, where humanity hands off control to transformative AI at some point. Hopefully it’ll be aligned with us and that’ll be a good transition that goes well, but it might not be. The stakes are incredibly high.

3: Build defences and diffuse AGI

This strategy argues for everyone to have access to their own AGI and compute, but for tremendous resources to be invested into defensive technologies. Proponents believe that the best future with AI is one where access and control over AI is widespread, and no single actor or group has too much power over this transformative technology.

Prominent proposals of this strategy are “d/acc” by Vitalik Buterin, “def/acc” by Matt Clifford, and to some degree “Differential Technological Development” by Nick Bostrom.

In the future this strategy envisions, protective technologies outpace destructive technologies. For example, even if millions of people could build viruses worse than the ones which caused the COVID pandemic, humanity would be fine because we’ve built rapid pathogen detection, containment and treatment capabilities. Even if every future teenager could build malware which could hack into a bank today, our financial systems are safe because the banks of the future have much more sophisticated cybersecurity. Future AIs would behave in desirable ways as a result of “AI alignment” techniques succeeding, like Anthropic’s Constitutional AI. And whatever new weapons of mass destruction are concocted, the world coordinates to build defences against those weapons before they’re able to cause catastrophic harm, and ideally before the weapon is ever built.

Critiques of this strategy focus on our inability to guarantee that we _can_ or _will_ build defences against new types of attacks. For example, Bostrom’s “vulnerable world hypothesis” proposes a thought experiment: what if we lived in a world where making nuclear weapons was easy, using materials available to everyone? What protective technologies could we build in such a world?

Another critique of this strategy is that while an attacker only needs to find one civilisational vulnerability to exploit, the defenders must protect and fix all exploitable vulnerabilities. It might be the case that the amount of resources needed to defend dwarfs the amount of resources needed for a bad actor with powerful AI to cause harm.

————————————

Note that some people believe we should pursue all three strategies, e.g. that we should slow down and control AI development in the short-term via government AGI projects, that we need to build defences against AI harms in the medium-term as a result of a diffusion of frontier capabilities, and that in the long-term we should relinquish control to an aligned superintelligence.

Unit 5 · Defence in Depth

Building defences

In the last unit, you created a threat scenario and a kill chain.

Read about each layer of defence in the next 3 sections, then you’ll apply ONE layer of defence to slow down or stop the attack.

A reminder of the layers: 1. Prevent the training of a dangerous AI model

E.g. stop an AI company from training their next generation of AI model

How could you prevent the training and creation of the AI system with the capabilities required to execute on your attack?

2. Constrain a dangerous AI model's capabilities and actions

E.g. an AI company ensuring that its new, dangerous AI model remains under their control and doesn't take dangerous actions in the world

If a dangerous AI is trained, how could an AI company constrain its ability to cause harm?

3. Withstand dangerous AI actions

E.g. beef up society's resilience to AI-enabled attacks via biosecurity, cybersecurity, and democratic hardening

If the dangerous AI systems escape all controls, how could we make society more resilient to its dangerous actions?

Exercise — Applying defences against your attack.

Return to your character from Unit 1, and your kill chain from Unit 4.

Which defensive layer do you think is most critical for your kill chain?

Now describe what it would look like for your character to implement that defensive layer, to protect humanity from your kill chain.

Unit 5 · Defence in Depth

Layer 1: Prevent dangerous AI training ≈30m

The first layer of defence is prevention: not training dangerous AI systems in the first place.

Interventions that (mostly) live here: export controls and the most advanced AI chips, international agreements and regulations not to train AI systems that cross certain red lines, “safe by design” AI systems, and norms in the research community not to race ahead in a reckless way.

Prevention feels the cleanest. If nothing dangerous gets built, nothing bad can happen. But it’s also one of the hardest layers to establish. There’s intense commercial competition between AI companies, and intense rivalry between nations. Most actors are incentivised to race ahead.

What would it take to prevent any actor from training an AI system with civilisation-threatening capabilities? How could we overcome tremendous geopolitical and commercial incentives to race ahead?

Leading AI scientists propose an international treaty that would cap the computing power used to train AI models and create a collaborative AI safety laboratory ("CERN for AI"). This is one model for how governments could coordinate to prevent the development of potentially catastrophic AI systems.
Misc · 5m
A former OpenAI researcher argues that private AI companies cannot safely develop superintelligence due to security vulnerabilities and competitive pressures that override safety. He argues that a government-led 'AGI Project' is inevitable and necessary to prevent adversaries stealing the AI systems, or losing human control over the technology.
Leopold Aschenbrenner · 25m
Unit 5 · Defence in Depth

Layer 2: Constrain dangerous AI capabilities ≈40m

When AI companies train their next generation of AI systems, they don’t know what that AI system will be capable of. They know that if they make it bigger and train it for longer, it will be better, but they don’t know how much better and in what way.

This also means that they don’t know how dangerous their next generation of AI systems will be. They discover that during and after the training process. Then they implement safeguards to try to prevent their AI systems from taking harmful actions.

So, if an AI company trains a dangerous AI system, how will they know that they’ve done so? And what might they do about it?

AI systems regularly do things their creators never intended, from maze-solving AIs that get stuck in corners to social media algorithms promoting extremist content. In this blog post, Adam Jones explains how outer alignment (setting correct goals) and inner alignment (ensuring AI follows those goals) might help to prevent these failures as systems become more powerful.
Adam Jones · 15m
This blog post explains why it might be easier to build walls around dangerous AI systems than to make them genuinely care about human welfare. However, these may only be temporary fixes before AI becomes too powerful to control.
Sarah Hastings-Woodhouse · 10m
In this report, RAND researchers identify real-world attack methods that malicious actors could use to steal AI model weights. They propose a five-level security framework that AI companies could implement to defend against different threats, from amateur hackers to nation-state operations.
Nevo et al. · 15m

Exercise — Learn how weak existing guardrails are.

Pliny the Liberator is an online persona who jailbreaks AI systems. They've created this database of prompts that you can use to bypass guardrails on AI models.

Try this out for yourself to understand how vulnerable these AI systems are to simple attacks.

Describe your process and results below.

A quick note before you begin:

Attempting jailbreaks on commercial AI platforms (ChatGPT, Claude, Gemini, etc.) may violate their Terms of Service and can result in account suspension. We recommend:

Using a secondary account you don't rely on for other work

Testing on open-source models locally (e.g., via Ollama)

Stopping if you receive any warning from the platform

This exercise is meant to be about understanding vulnerabilities not about encouraging ToS violations.

Unit 5 · Defence in Depth

Layer 3: Withstand dangerous AI actions ≈30m

What happens if a dangerous AI model is trained, and it bypasses an AI company’s safeguards, or escapes their control entirely?

If dangerous AI models roam wild, we need to harden society against a possible enslaught of attacks.

Imagine what it would look like to have a pandemic-proof, where no matter how good an AI system is at building bioweapons, we detect and contain pathogens immediately.

Imagine what it would look like for critical national infrastructure to be secure against physical and cyber attacks. We might have the most capable attackers working “on our side”, testing the defences, highlighting where there are weaknesses, and using their expertise to shore up the defences.

Imagine what it would look like for our democratic systems to elect wise, benevolent leaders and for our information ecosystems to enable people to make sense of reality together.

Ethereum founder Vitalik Buterin describes how democratic, defensive and decentralised technologies could distribute AI's power across society rather than concentrating it, offering a middle path between unchecked technical acceleration and authoritarian control.
Vitalik Buterin · 30m

Optional resources
Read the abstract and executive summary (pages 1 to 7). Yoshua Bengio (the world's most cited computer scientist) and other leading AI researchers describe how today's AI training methods systematically produce dangerous behaviours. They then propose 'Scientist AI' as an alternative path: non-agentic AI systems designed to understand rather than act.
Bengio et al. · 15m
Sophie-Charlotte Fischer, Jade Leung and Markus Anderljung et al.
Sarah Hastings-Woodhouse
An introduction to mechanistic interpretability (techniques for understanding AI internal reasoning) that shows how researchers could detect when models are deceiving users or cutting corners to achieve goals.
Sarah Hastings-Woodhouse · 10m
You can't effectively constrain dangerous AI capabilities without first detecting and measuring them. CSET researchers explain the difference between testing what AI models can output and how they affect real-world outcomes, and how this is used to measure risk.
Jessica Ji, Vikram Venkatram, and Steph Batalis · 10m
Leopold Aschenbrenner · 25m
Markets excel at safety when the risks hit consumers directly (e.g. unsafe aircraft), but fail with long-term, diffuse risks like climate change and AI threats. In this blog post, Nielsen proposes institutional changes that might accelerate defensive technologies faster than offensive ones.
Michael Nielsen · 40m
Read the Abstract, Policy Recommendations, and Introduction. This paper uses scenario planning to show how governments could prepare for AI emergencies. The authors examine three plausible disasters: 1) losing control of AI, 2) AI model theft, and 3) bioweapon creation. They then expose gaps in current preparedness systems, and propose specific government reforms including embedding auditors inside AI companies and creating emergency response units.
Wasil et al. · 10m
Jamie Bernardi argues that we can't rely solely on model safeguards to ensure AI safety. Instead, he proposes "AI resilience": building society's capacity to detect misuse, defend against harmful AI applications, and reduce the damage caused when dangerous AI capabilities spread beyond a government or company's control.
Jamie Bernardi · 10m
Read Section 6: Breaking the Intelligence Curse
Luke Drago and Rudolf Laine
Unit 6 · Start Contributing

Choose your focus ≈20m

Core ≈ 2h30

In the last unit, we explored what defences need to be built in order to defend against AI harms.

Very few people in the world are trying to steer the development of AGI towards beneficial outcomes for humanity. We estimate it’s less than 2,000 full-time people!

You could be one of the few people making a difference, but doing so requires stepping up, taking responsibility, and trying to solve new problems that don’t have an instruction manual.

In this unit, you'll create an action plan for how you could start contributing. This is only the beginning - you're not expected to have a perfect answer right now.

You now have an opportunity to take ownership and _just start making AI go better_. We're really excited about helping you on this journey, and to see what you do! A common failure mode is to think "Oh, I can't actually do X" or to say "Someone else is probably doing Y." You probably can do X, and it's unlikely anyone is doing Y! It could be you!
Neel Nanda · 5m
Read this long list of interventions, which are organised by defensive layer. This gives you a flavour for the kind of interventions that could be pursued to make AI go well. Consider what else you learnt during the course and where you think important work is needed - this list is not exhaustive!
Dewi Erwan · 15m

Exercise — Prioritise a single intervention.

Based on everything you've learnt about so far during this course, prioritise an intervention from the long list.

To help you prioritise, consider which intervention you think would be effective against the kill chain you developed in Unit 4.

Unit 6 · Start Contributing

Go deep in one area

Exercise — Do your own research.

- What does success look like with this intervention? How does it help make AI go well?

What's the current status of this intervention? Are governments, AI companies or other actors already doing it? If not, why not?

Which organisations are working on this which you can contribute to or join?

Spend ~1 hour on this.

Unit 6 · Start Contributing

Make a plan ≈10m

In this unit, you'll synthesise everything you've learnt so far during the course, and you'll develop your own action plan for how you could contribute to making AI go well.

We don't expect this to be super detailed or a perfect plan - treat this as a first draft, and something that's _good enough_ to help you get started.

Use this template to develop your own personal action plan. If you think a different template would work better for you, we encourage you to deviate!
Dewi Erwan · 10m

Exercise — Write your personal action plan.

Spend ~1 hour on this.

Once you're finished, 1) make the google doc shareable, 2) paste it below, and 3) share it with your group in Slack!

Unit 6 · Start Contributing

Create your 1-pager (optional)

Now that you've explored the landscape of AI safety, it's time to clarify your own path forward.

Start by creating your 1-pager — a concise document that captures what you're looking for and what you bring to the field. This is useful both for your own thinking and as something you can share with potential employers or collaborators.

We don't expect this to be super detailed or a perfect representation. Treat this as a first draft, and something that's _good enough_ to help you get started.

Use this template to clarify what you're looking for and showcase what you bring to AI safety. It's useful for helping you think through your own positioning, and giving you something concise to share for future opportunities. If a different format works better for you, feel free to deviate!
BlueDot Impact

Exercise — Create your 1-pager.

Spend ~1 hour on this.

Once you're finished, make the Google Doc shareable and share it with your group in Slack!

Unit 6 · Start Contributing

Optional advice ≈2h00

Check out these resources if you'd like to receive career advice on AI governance, technical AI safety, or biosecurity.

Richard Ngo, a former researcher at OpenAI and Google DeepMind, argues that finding your personal fit matters far more than choosing the "right" AGI safety career path, and that the field's biggest bottleneck is people with agency who actually implement concrete proposals rather than abstract ideas.
Richard Ngo · 20m
80,000 Hours explains how to contribute to AI policy—through government roles, think tank research, or tech companies—with concrete guidance on building relevant skills, choosing between technical and policy paths, and finding your first opportunities.
Cody Fenwick · 50m
80,000 Hours breaks down AI safety research careers into achievable steps—showing that strong software engineers can contribute within a year while theoretical roles typically need PhDs—and provides concrete guidance on learning paths, testing your fit, and where to apply.
80000 Hours · 50m
Unit 6 · Start Contributing

Next steps: Deeper-dive courses

This course gave you the strategic landscape: drivers of AI progress, threat pathways, and layers of defence.

Our deeper-dive courses let you specialise in the area that fits your background and where you want to contribute. Intensive and part-time versions of each of these courses kick off every month.

For people with technical backgrounds considering policy, and policy professionals adding AI expertise. Technical safety work alone won't be enough. Governance is one of the key levers for ensuring AI goes well. This course covers how decisions get made, who has power, and which agencies matter. You'll develop fluency in major proposals: compute governance, safety standards, liability frameworks, international coordination.
BlueDot
For engineers, scientists, policy professionals, and entrepreneurs who want to defend against pandemic threats. The AGI Strategy course introduced catastrophic pandemics as one of the key threat pathways. This course goes deep on the defences: preventing engineered outbreaks, detecting early infections, and deploying rapid countermeasures. You'll explore how AI and synthetic biology are changing the threat landscape, and what it would take to build a pandemic-proof world.
BlueDot
For ML researchers, software engineers, and policy professionals looking for technical depth. This course takes the "layers of defence" framework from AGI Strategy and asks what it takes to implement each layer. You'll learn current safety techniques across training, deployment, and monitoring, and develop your own view on which approaches are most promising.
BlueDot
Unit 6 · Start Contributing

Next steps: Programs