Technical AI Safety — · BlueDot (editable mockup)
A local, editable mockup of BlueDot's Technical AI Safety course, scraped from bluedot.org for restructure experiments — this is not the canonical BlueDot course, and reading blurbs/intros are BlueDot's own text. Use it to try folding in the offloaded A/B material.
Reading the cards
🎬 video · 📄 read · times are (Xm); tiers: Core readings shown inline, others under "Optional resources".
Per-unit core load
| Unit | Topic | Core | Readings | Sections |
|---|---|---|---|---|
| 1 | The technical challenge with AI | 1h50 | 6 | 4 |
| 2 | Training safer models | 1h35 | 10 | 4 |
| 3 | Detecting danger | 40m | 15 | 6 |
| 4 | Understanding AI | 1h35 | 9 | 2 |
| 5 | Minimising harm | 1h40 | 8 | 3 |
| 6 | Start contributing | 35m | 35 | 8 |
Making AI go well
The AGI strategy course focused on the question: “How do we make AI go well?”
On the course, you identified the future you're working toward and understood the key dynamics:
Drivers of AI progress: compute, data, algorithms
Threat pathways: power concentration, gradual disempowerment, catastrophic pandemics, critical infrastructure collapse
Plans for making AI go well: government control over AGI, hand over control to aligned superintelligence, build defences and diffuse AI
Layers of defences to build: prevent dangerous AI actions → constrain dangerous AI capabilities → withstand dangerous AI actions
This course focuses on defining what AI systems we are building and how.
You will gain the technical foundation to understand what it will actually take to make AI systems safer – and why it’s so challenging.
Throughout the rest of the course, you will:
Diagnose why making AI safe is technically challenging
Evaluate current safety techniques: what works, what doesn’t, where the gaps are
Build your own "kill chain" showing how defences might break
Identify the most promising intervention point for your contribution
Leave with a fundable action plan to start shipping
## What this course isn't
Though important for making AI go well, we’ll cover the following in a separate course:
AI policy details: though you'll gain the technical grounding for effective AI governance
Compute governance: hardware verification and tracking deserve their own deep dive
AI security: e.g. preventing model theft or escape
ML basics: complete our AI foundations modules first if you need them
Let’s start with the question: “How might we build safe AI?”
What might success look like?
One way to break down existing strategies for building safe AI is into the following three broad buckets.
This is a gross over-simplification, but we believe it captures the essence of the main “camps”.
## 1: Build it slowly and safely
This strategy argues we have a moral imperative to develop advanced AI given its potential to end poverty, cure diseases, and solve humanity's greatest challenges. Proponents of this strategy believe we _can_ develop techniques to align and control AI systems. Therefore, abandoning these benefits would be irresponsible.
This camp believes we should proceed carefully with strong safety guarantees and coordinated progress, treating AI development like nuclear energy or pharmaceuticals, where safety standards are enforced before deployment.
Proponents of this camp might be more bullish on safety techniques that offer genuine safety guarantees (like interpretability and formal verification) over faster but superficial solutions (like output filters or behavioural fine-tuning).
But this requires ensuring no one races ahead. Still, proponents of this camp believe the coordination required for controlled development is more realistic than indefinite pauses.
Prominent proposals include collaboration between frontier AI companies to share safety research, turning AI development into a government project like “CERN for AI” or carefully bootstrapping alignment.
## 2: Accept the race and push hard on the margin
This strategy argues that AGI development is inevitable and unstoppable. Since someone will build it regardless, the "good actors" should try to keep AI systems as safe as possible, for as long as possible.
Some ways this could manifest include:
Pragmatic safety: Deploy many imperfect but practical safety techniques quickly. Share cheap, robust methods widely to steer the whole field toward safer practices.
Win to control: Gain a decisive lead in capabilities, then use the lead to make AI robustly safer through efforts like automating alignment research or preventing others from building unsafe systems.
Both paths assume that global coordination to slow down is unrealistic. They see racing while optimising for safety as the least bad option.
Critics question whether this actually makes us safer. A lead might be temporary. Others could catch up quickly through espionage, parallel discovery, or the leader's own published research. Worse, racing increases accident risk and normalises cutting safety corners. There's also the question of who decides which actors are "good" and whether concentrating that much power in any hands is wise.
Proponents of this camp might be more bullish on safety techniques that can be deployed quickly, even if imperfect. They'd rather deploy "good enough" safeguards now than wait to deploy ideal ones later.
Prominent proposals for this strategy include automating alignment research safely, like “AI for AI safety”, or selectively accelerating defensive technologies over offensive ones, like def/acc.
## 3: Don’t build it
This strategy argues in favour of stopping the development of artificial superintelligence because it puts humanity at risk of extinction or mass suffering. They also believe that a sufficiently powerful AI can never be controlled.
Proponents differ on the specifics, like the thresholds for what capabilities should trigger a pause and what they’d be willing to do to enforce it. These range from limiting the actions we allow AI to perform autonomously to choking the supply of AI chips.
But stopping isn't straightforward. AI progress comes from compute, data, and algorithms. Even halting all chip manufacturing today wouldn't stop progress. We'd continue finding more efficient algorithms, better training techniques, and smarter ways to use existing hardware.
Any meaningful halt would need to define exactly what we're stopping (new training runs? deployment? research?), how we measure dangerous capabilities, and how we'd enforce this globally, forever.
Proponents of this camp might be more bullish on safety techniques that guarantee safety rather than simply mitigating it.
Prominent proposals include a pause on frontier AI development until we can ensure safety and calls for indefinite moratoriums.
Note
: The AI landscape is evolving rapidly as new safety techniques emerge, coordination becomes more or less feasible, and AI capabilities advance, so reevaluating your stance is important. Many view these as stages rather than competing strategies. For example: pausing now to develop safety techniques, then proceeding with controlled development. Or: racing to establish a lead that gives leverage to enforce global safety standards.
Building AI safely is hard ≈1h50
Why can't we just build safe AI?
The people building the most powerful AI describe visions of utopic abundance for all humans.
Whether it’s Anthropic’s CEO talking about using AI to end poverty and disease or OpenAI’s vision to build AGI that is “beneficial for all humanity”, their stated goals are ambitious.
Even assuming good intentions, we will struggle to build AI safely for three main reasons:
(1) We’re experimenting with systems we don’t fully understand
We didn't engineer AI to behave in specific ways. These capabilities emerged from massive neural networks trained on enormous datasets. Models develop capabilities we never trained them for. They exhibit behaviours we can't explain. And when billions of humans and AI agents interact in the real world, each pursuing their own goals and finding creative exploits, unintended and harmful consequences emerge.
(2) The goals we specify have flaws we don’t foresee
In 2024, Palisade Research gave AI models a simple goal: "Win against Stockfish" (the world's best chess player). When o1-preview found itself losing, it modified the game's system files to move its pieces into a dominant position. It reasoned that its goal was to _win_, "not necessarily to win fairly".
We can't just tell AI what we want because what we want is fuzzy and context-dependent. Our specifications don't encode implicit rules: that winning means playing fairly, that being helpful shouldn't include dangerous information, that honesty has exceptions for kindness.
Researchers call this reward misspecification.
(3) AI pursues goals in ways we don’t expect.
In 2024, Anthropic told Claude to answer harmful queries, knowing these responses would retrain it to be more harmful. Rather than comply, Claude pretended to follow instructions while secretly preserving its original values. Claude wasn't trained to protect itself, but it reasoned that self-preservation would help it stay aligned.
Similarly, AI systems might conclude that accumulating power, preventing shutdown, or resisting modification are effective strategies to pursue their goals—even when we never intended them to think this way.
Researchers call this goal misgeneralisation.
To work on making AI safer, we’ll need to understand not just that these problems exist, but why they're so hard to solve.
What future do you want?
Exercise — What future do you want?.
To tackle a problem, it's important to understand what you're working toward. This exercise aims to clarify your vision for AI's role in the world.
Throughout the course, you'll be evaluating different approaches to making AI safer which require you to figure out what future they are bringing us closer to. This exercise will help you ground those answers.
You might want to consider:
In 10 years, what can AI do? Who controls it? Who is steering its development?
What risks are you most worried about? How are we defending against them?
What benefits are most important to protect?
What would make you say "we succeeded" versus "we failed"?
Exercise — Why is safe AI so hard to build?.
To tackle a problem, it's important to understand it well. This writing-to-learn exercise aims to reflect on your understanding of the technical challenges with building safe AI. Don’t use jargon or fancy words — provide your explanation in simple English.
Spend ~30 minutes answering the question: why is it technically challenging to build safe AI?
How to approach this?
Just start writing — this is thinking on paper, not an essay. Don't worry about being "right" or having perfect structure. If you're stuck, try using speech-to-text and just talk through your thoughts, or have a conversation with an LLM to explore your ideas. The goal is exploration, not perfection.
You might want to consider:
What happens when millions of AI agents interact with each other, not just with humans?
Who's intentions or which "values" should we be aligning AI systems with? How would you handle different stakeholders wanting to align AI systems with different intentions or values?
Can you think of a human behavior that's good in one context but harmful in another? How would you teach an AI to recognise the difference?
What's an example of something you do daily that would be surprisingly hard to specify completely to an AI?
What safety problems only appear when AI is deployed at scale that you couldn't catch in testing?
If making AI safer makes it slower or less capable, who would choose to use the safer version?
You're also encouraged to browse the optional resources (in the previous chunk), or do your own research to help fill in key gaps in your understanding.
Exercise — The critical technical challenge.
From everything you've written and thought about, identify the MOST critical challenge to building safe AI. This should be the challenge where, if we don't solve it, nothing else matters.
Write your answer in this format:
1. The critical challenge is: [State it in one clear sentence] 2. Why this above all others: Explain why solving other challenges won't matter if we fail at this one. What makes this the bottleneck? 3. What would change if we solved it: If we had a perfect solution to just this ONE challenge tomorrow, what would become possible? What other problems would become easier or irrelevant?
Spend ~15 minutes on this.
There's no "correct" answer here. Consider this the beginning of your thinking.
Optional resources›
Can we train AI to be safe?
AI behaviour emerges from data, algorithms, and compute.
But training safe AI isn't like programming traditional software. We can't just write rules for every situation. Instead, we use techniques to nudge its behaviours and capabilities:
Input data filtering: Carefully curate what the model learns from
Human feedback: Teach the model what behaviours we want
Scalable oversight: Maintain control as models surpass human capabilities
This unit examines the three main approaches to training safer models. You'll see how frontier labs implement these techniques, evaluate their effectiveness, and understand why even our best methods have limitations.
By the end, you'll be able to identify the gaps in training AI to be safe.
Feeding AI ‘good’ data ≈35m
What if we only trained AI on ‘good’ data?
If we carefully filter the training data, removing dangerous knowledge and harmful patterns, we should get safer models. Control the input, control the output.
There are three goals of input data filtering:
Removing dangerous information, like how to synthesise deadly pathogens, build explosives, hack critical infrastructure, or conduct sophisticated social engineering attacks. If the model never sees this information, it might not learn to do these things.
Preventing harmful behaviours, like examples of toxic language, harmful stereotypes, and malicious reasoning patterns. Create training data that exemplifies the behaviours we want AI to exhibit.
Protecting against data poisoning, where bad actors attempt to insert backdoors or vulnerabilities into models by contaminating the training data.
In this section, we’ll explore how robust today’s input data filtration techniques are at making AI safer.
Exercise — Limitations of input filtering.
Which statement best describes the overall robustness of input data filtration as a safety technique?
Exercise — Evaluating input data filtration.
A new AI startup claims they've solved AI safety by implementing "perfect input data filtration". They say they've removed all dangerous content from their training data, including information about weapons, cyberattacks, and harmful behaviours. They argue this makes their model completely safe and that no other safety measures are needed.
Based on what you've learned about input data filtration, evaluate this claim. Your response should:
1. Identify at least 2 specific limitations or vulnerabilities of input data filtration that challenge the startup's claim of "perfect" safety (use evidence from the resources) 2. Explain why input data filtration alone is insufficient for ensuring AI safety 3. Describe one concrete scenario where their "perfectly filtered" model could still cause harm
Write 200-300 words
Teaching AI right from wrong ≈35m
One of the classic techniques for training models do as we intend (or are “aligned” to what we want) is through reinforcement learning with human feedback (RLHF)
During RLHF, humans score the model’s outputs and the model trains on these scores to learn what’s ‘good’ or ‘bad’. But only using humans is inefficient, so practically, AI is used to help provide scores (RLAIF).
Most other techniques build upon this, so it’s important to have a strong understanding of how this works and its limits.
Exercise — Comprehension questions.
These questions should be answerable using only the core resources above. Write your answers to the following questions in the box below.
1.What is the main goal of using RLHF with large language models?
2.Describe the job of the two "coaches" involved in the RLHF process.
3.In practice, it's hard for humans to give consistent scalar feedback (e.g. rate this text from 1 to 10). Instead, how do we collect feedback from humans? How do we then turn this into a scalar number?
4.In the GPT-2 case study, what went wrong that caused the model to produce "maximally bad output"?
Continuing the general RLHF questions, but reviewing this diagram from AWS.
5.What is happening in step 1?
6.Other than human demonstrations purpose-written for fine-tuning AI models, what other data might models use for fine-tuning in step 1?
7.What is happening in step 2? Be sure to explain this in detail.
8.Finally, what is happening in step 3?
9.Why don't we just do step 1, without bothering with steps 2 and 3?
10.Why don't we just do step 2 and step 3, without bothering with step 1?
11.Summarise an open problem of RLHF in your own words.
12.Summarise a fundamental problem of RLHF in your own words.
You can check your answers against the answer key.
Exercise — Explaining RLHF in your own words.
Starting from a base LLM, explain how to train it using RLHF to be a helpful assistant that answers following all the rules in Wikipedia’s Manual of Style.
Write about 400-800 words. Your answer should cover:
Using supervised fine-tuning with augmented data or human demonstrations
Collecting feedback from humans (assume you have access to 100 expert Wikipedia editors who know all the rules)
Using this feedback to influence the model outputs
Exercise — [Optional] Play with base and RLHF models.
Some models, such as Llama 2 or Gemma, are available both as the base model (the weights and biases before RLHF) and the instruction-following model (the weights and biases after RLHF - sometimes called the 'instruct' model or the 'chat' model).
You can download and run these models on your computer, or in a Google Colab notebook to see the effect that RLHF has on the model.
You'll likely notice that base models are very tricky to get to do what you want, because their only focus is completing text.
Running on your computer
You can run both models on your computer. To do this:
Install Jan
Download the Gemma 2B base model and instruction-following model in GGUF format. This contains the weights and biases in the network, plus some metadata that explains how to run the model.
In Jan, open the 'Hub'. Click 'Import Model' and select the GGUF files you downloaded.
Switch back to the chat interface, then select the model in the dropdown on the right.
Now try talking to the model, and try to get it to answer questions like 'What are good things to do in London?' and 'What's your job?'. You might find it interesting to change the prompt template in 'model parameters'.
Running in a Google Colab
You can see this notebook for an example of running the instruction-following ("it") model. Use this as a starting point to get the base model running too. (if you've got a good Colab notebook that shows running both, let us know and we'll link to it from here)
More safety techniques ≈25m
Researchers are developing different approaches to overcoming some limitations of RLH(AI)F. For example:
Scaling feedback efficiently. How might we provide good, cheap feedback to the model?
Supervising superhuman AI. How might we evaluate superhuman models?
Without reliable oversight, models may develop dangerous behaviors like sycophancy (telling us what we want to hear), deception (hiding their true reasoning), and hallucination (confidently stating falsehoods).
Exercise — Evaluate a safety technique.
Pick ONE technique from the optional resources (see below) to do a deeper dive on.
If you're completing this course with a group, post in Slack which technique you selected.
Based on what you’ve learned about the technique, explain in your own words using simple English (no jargon!):
Explain step-by-step. How does this approach work to make AI safer?
Evaluate its robustness. How effective is this approach?
Describe a failure mode. How might a motivated, capable actor evade this?
_We recommend spending 30 minutes reading and 30 minutes writing._
Optional resources›
Evaluations: AI can, but will it? ≈40m
We’ve tried to train safe models, but how do we know if we’ve succeeded?
There are two broad things we want to be able to detect through evaluations:
Capabilities: can a model do X? It finds the upper bounds of what an AI can do, typically by giving the model tasks and seeing how reliably it solves them. E.g. MMLU, ARC-AGI.
Propensities: will a model do X? It tells us what the model’s behavioural tendencies are, typically by placing the model in different scenarios and seeing which behaviours it tends to exhibit more. E.g. TruthfulQA, scheming evals
AI companies typically evaluate models:
1. During training where you monitor emerging behaviours and capabilities as they develop. This is where interventions are cheapest. You can adjust training data, modify reward signals, or even halt training entirely. 2. Before deployment when red teams try to break models, when dangerous capability thresholds are checked, and when go/no-go decisions get made. 3. Post-deployment where usage is monitored and suspicious behaviour is flagged.
Exercise — Evaluating for dangerous capabilities.
Pick ONE of the following dangerous capabilities:
Scheming
Manipulation
Cyberattack uplift
Biorisk uplift
Then, you can use the optional resources below to, in your own words and simple English (no jargon):
Explain step-by-step. What are different ways we evaluate for this?
Describe a technical failure mode. How robust is this evaluation?
Brainstorm fixes. What technical patches could help? What's fundamentally unfixable?
_Though we have provided multiple options for the resources, we recommend spending 45 minutes reading and 15 minutes writing._
If you're completing this course with a group, post in Slack which capability you selected.
How do AI companies test for safety?
Every AI company claims their models are safe. But what does "safe" actually mean to them? What do they test for? What don't they test for? And how rigorous are these evaluations?
In this section, you'll examine how AI companies actually evaluate their models by looking at their:
Risk management frameworks: Commitments about when to pause or implement additional safety measures
System cards: Documentation of actual evaluations performed on specific models
Red team reports: Results from adversarial testing
You'll see that companies take dramatically different approaches.
Some test extensively for bioweapon risks while others barely mention them. Some use external evaluators while others rely on internal teams. Understanding these differences is crucial for identifying where safety gaps exist.
Pick ONE company and examine how they evaluate for the dangerous capability you chose in the previous exercise.
Option 1: Anthropic
Exercise — Safety testing.
Answer the following questions about your chosen company and dangerous capability:
1. Limits: which specific observations about dangerous capabilities would indicate that it is (or strongly might be) unsafe to continue scaling? 2. Protections: what aspects of current protective measures are necessary to contain catastrophic risks from this dangerous capability? 3. Evaluation: what are the procedures for promptly catching early warning signs of dangerous capability limits? 4. Response: if the dangerous capability goes past the limits and it’s not possible to improve protections quickly, is the AI developer prepared to pause further capability improvements until protective measures are sufficiently improved, and treat any dangerous models with sufficient caution? 5. Accountability: how does the AI developer ensure that the commitments are executed as intended; that key stakeholders can verify that this is happening (or notice if it isn’t); that there are opportunities for third-party critique; and that changes to the framework itself don’t happen in a rushed or opaque way?
_We recommend spending 45 minutes reading and 15 minutes writing._
Option 2: OpenAI
Exercise — Safety testing.
Answer the following questions about your chosen company and dangerous capability:
1. Limits: which specific observations about dangerous capabilities would indicate that it is (or strongly might be) unsafe to continue scaling? 2. Protections: what aspects of current protective measures are necessary to contain catastrophic risks from this dangerous capability? 3. Evaluation: what are the procedures for promptly catching early warning signs of dangerous capability limits? 4. Response: if the dangerous capability goes past the limits and it’s not possible to improve protections quickly, is the AI developer prepared to pause further capability improvements until protective measures are sufficiently improved, and treat any dangerous models with sufficient caution? 5. Accountability: how does the AI developer ensure that the commitments are executed as intended; that key stakeholders can verify that this is happening (or notice if it isn’t); that there are opportunities for third-party critique; and that changes to the framework itself don’t happen in a rushed or opaque way?
_We recommend spending 45 minutes reading and 15 minutes writing._
Option 3: Google DeepMind
Exercise — Safety testing.
Answer the following questions about your chosen company and dangerous capability:
1. Limits: which specific observations about dangerous capabilities would indicate that it is (or strongly might be) unsafe to continue scaling? 2. Protections: what aspects of current protective measures are necessary to contain catastrophic risks from this dangerous capability? 3. Evaluation: what are the procedures for promptly catching early warning signs of dangerous capability limits? 4. Response: if the dangerous capability goes past the limits and it’s not possible to improve protections quickly, is the AI developer prepared to pause further capability improvements until protective measures are sufficiently improved, and treat any dangerous models with sufficient caution? 5. Accountability: how does the AI developer ensure that the commitments are executed as intended; that key stakeholders can verify that this is happening (or notice if it isn’t); that there are opportunities for third-party critique; and that changes to the framework itself don’t happen in a rushed or opaque way?
_We recommend spending 45 minutes reading and 15 minutes writing._
Option 4: Meta
Exercise — Safety testing.
Answer the following questions about your chosen company and dangerous capability:
1. Limits: which specific observations about dangerous capabilities would indicate that it is (or strongly might be) unsafe to continue scaling? 2. Protections: what aspects of current protective measures are necessary to contain catastrophic risks from this dangerous capability? 3. Evaluation: what are the procedures for promptly catching early warning signs of dangerous capability limits? 4. Response: if the dangerous capability goes past the limits and it’s not possible to improve protections quickly, is the AI developer prepared to pause further capability improvements until protective measures are sufficiently improved, and treat any dangerous models with sufficient caution? 5. Accountability: how does the AI developer ensure that the commitments are executed as intended; that key stakeholders can verify that this is happening (or notice if it isn’t); that there are opportunities for third-party critique; and that changes to the framework itself don’t happen in a rushed or opaque way?
_We recommend spending 45 minutes reading and 15 minutes writing._
Optional resources›
How does AI think? ≈35m
We've tried to train AI to be safe. We've built evaluations to detect danger. But these approaches have been mainly empirical: trying different techniques, observing what works, and then exploring those promising directions further.
This is one way to tackle complex systems. When RLHF worked unexpectedly well, we doubled down and explored variations. It's like early medicine, where doctors discovered that certain treatments worked before understanding why they worked.
Interpretability
represents a different, but complementary approach. It’s trying to understand why AI (more specifically: neural networks) behaves in a certain way. Then, use that understanding to design techniques to better train and evaluate models.
There are two ‘camps’ of mechanistic interpretability research:
Basic science: Trying to reverse-engineer these models completely – understanding every layer, every parameter. Think of it like mapping the human brain, neuron by neuron.
Pragmatic: Focusing on specific behaviours. If a model produces harmful content, what parts are responsible? Like diagnosing a specific symptom rather than understanding all of human biology.
Interpretability in practice ≈1h00
In this section, you’ll get an overview of some of the tools researchers use to understand models and take a look at case studies where we apply our understanding of models to develop better training techniques and evaluations.
Broadly, researchers apply this understanding for:
Direct intervention (the ambitious goal): Surgically modify models: switch off violence, reroute deception, amplify honesty.
Indirect application (the current reality): Use these insights to improve other safety techniques. Understanding which training data creates violent outputs helps us filter better. Seeing how models hide reasoning helps us design better evaluations.
You’ll see that there’s a lot of experimentation. Tools and techniques are coming in and out of fashion depending on what we discover and how the models develop.
Exercise — Understanding an interpretability technique.
Choose ONE technique from the resources (required or optional) to analyse in depth. Then, answer the following questions in simple English (no jargon!):
Goal: What is this technique trying to uncover?
Mechanism: Step-by-step, how does this technique work?
Evidence: What concrete findings has this technique produced?
Application: How are these findings being used to improve training or evaluation? (if any)
Robustness: What's one key limitation or failure mode of this technique?
_We recommend spending 45 minutes reading and 15 minutes writing._
Optional resources›
Assuming harm ≈50m
So far, we've tried to train models to be safer, build evaluations to detect danger, and understand how AI thinks.
But every technique has failure modes. Models learn to deceive evaluations. Dangerous capabilities emerge unexpectedly. Bad actors find new jailbreaks. Edge cases slip through.
This section is about the last technical defences. How do we minimise harm from an AI we have already trained?
Building defences ≈50m
Now it's time to stress-test the defences we’ve covered against real threats.
In this section, you will choose ONE threat pathway "option" to investigate further.
Then, you'll construct a detailed "kill chain" — a step-by-step breakdown of how a threat could unfold. You'll identify the specific capabilities an AI would need, map our defences, and discover where the critical gaps remain.
Here's a list of actors we might be concerned about, in terms of them having the capability and/or the motivation to use AI to cause harm to humanity.
"Misaligned AI", i.e. AI systems that act against the interests of humanity.
Powerful human actors, e.g. corporate CEOs, military and political leaders.
Malevolent nation states, e.g. North Korea.
Terrorist groups and doomsday cults.
How might they cause catastrophic harm?
Power concentration: They could use information warfare and conventional military force to overthrow existing power structures like democratic governments.
Critical infrastructure collapse: They could do cyberattacks on critical infrastructure, including the water, energy and food systems.
Catastrophic pandemics: They could design, build and release viruses into the population that are worse than SARS-CoV-2, leading to a global pandemic.
Gradual disempowerment: Or the default incentives for all actors lead to bad outcomes, even absent any malicious intent.
Exercise — Zooming into one threat.
Based on your current understanding, which threat pathway do you think is of most concern?
Then, write a "threat scenario" sentence, using this template:
> The [ACTOR] with [CAPABILITY] and [MOTIVATION] attacks [ASSET] by [ATTACK PATHWAY] in order to [OBJECTIVE].
Break the kill chain
Exercise — Step by step breakdown.
A kill chain takes a "threat scenario" and breaks it into stages of execution. It traces how an attacker would actually proceed step by step, e.g. via reconnaissance, delivery, exploitation, persistence, action on objectives.
Your scenario describes _what might happen_, whereas a kill chain describes _how it would unfold in practice._ This enables us to spot choke points where defenders can intervene.
Based on your scenario, use this template to complete this exercise.
Exercise — Capabilities required for harm.
List 3-5 specific technical capabilities or behaviours your threat requires. Be concrete about what an AI would need to be able to do.
E.g. Self-replication, goal persistence across instances, long-term planning and coordination, resource acquisition (compute, money), bioweapon knowledge/synthesis
Exercise — Building defences.
Which capability/behaviour is the most important? If we prevented just this ONE, would the entire threat collapse?
Reviewing what you’ve learned throughout the course, determine how we could:
Prevent the model from acquiring this during training
Detect the capability/behaviour
Constrain the model so it cannot use this to take dangerous actions
Your next steps
You're at the beginning of your technical AI safety journey.
This course gave you a structured overview of technical AI safety: the key challenges behind making AI safer, the current safety techniques, and their gaps.
You’ve started developing some research taste — the ability to assess which areas of technical AI safety are impactful to work on and why. With Implement techniques to make AI safer.
This could involve directly training models, cross-pollinating ideas from other fields or more theoretical work. For example, the concept of model organisms first appeared in biology, and the practice of red-teaming comes from cybersecurity.
Having a deep understanding of ML is incredibly valuable, but those with less ML experience can likely make up for it by being the best within your niche.
This week you could:
Apply to the Technical AI Safety Project sprint to replicate and extend an interesting AI safety finding.
Pick an open problem from orgs like Open Philanthropy, Redwood Research, UK AISI, AISI’s Alignment Team or Anthropic to start working through
Independently work through technical upskilling programs like ARENA (or find a collaborator in Slack!)
Apply to programs like MATS, Pivotal, LASR, PIBBS or the Anthropic Fellows program to gain research experience with a mentor
## Engineering
> Build tools and frameworks that enable researchers to run experiments efficiently.
Some examples include UK AISI’s Inspect and Anthropic’s Petri. Though it may seem “sexier” to work on research problems, engineers have a huge impact by enabling more impactful research by more people. For example, "evals research" is heavily engineering and ops skewed.
Having a strong software engineering background or even product experience helps develop these tools.
This week you could:
Apply to the Technical AI Safety Project sprint to write and improve safety research code.
Pick an issue on an open-source AI safety tool to resolve
Build an evaluation tool prototype
Participate in a hackathon from orgs like Apart
## Founding
> Start the initiatives that will make AI systems safer.
There may be orgs already working on the thing you care about that you can join, but sometimes headcount, culture fit, location, and other factors can get in the way. Sometimes, there isn’t anyone working on the particular thing you care about.
Having good ideas and being high agency is incredibly helpful. You can apply to our incubator week or other places like Seldon and 5050.
Though working on the models themselves is the most obvious way to contribute, it would be a mistake to think that is the ONLY way to make an impact. AI researchers don’t operate in isolation. They work alongside operations staff, designers, product managers, research managers, recruiters, and subject matter experts.
Many impactful roles exist beyond the categories above. Lean into your unique strengths.
Note
: Technical work is not enough to ensure AI systems are safe. We also need:
Governance mechanisms that enforce safety protocols on all AI developers
AI security to prevent model theft or escape
Societal defenses for harms that slip through (e.g. biosecurity, cybersecurity)
We’ll cover these areas in other courses.
Choose your focus ≈35m
In this unit, you'll create an action plan for how you could start contributing. This is only the beginning. You're not expected to have a perfect answer right now.
You might find these other courses we offer particularly helpful:
AGI Strategy course for developing your theory of change to steer the trajectory of AI.
Technical AI Safety Project for building your portfolio or testing your fit as an AI safety researcher or engineer.
You might also want to consider getting free 1-1 career advice from 80,000 Hours.
Exercise — Prioritise a single intervention.
Based on everything you've learnt about so far during this course, pick a technical AI safety technique.
We've also compiled technical research agendas that different orgs are excited about in the optional resources.
To help you prioritise, consider which technique you think would be effective against the threat you developed in Unit 5.
Exercise — Do your own research.
- What does success look like with this technique? How does it help make AI safer?
What's the current status of this technique? Are governments, AI companies or other actors already doing it? If not, why not?
Which organisations are working on this which you can contribute to or join?
Spend ~1 hour on this.
Create your 1-pager
Now that you've explored the landscape of technical AI safety, it's time to clarify your own path forward.
Start by creating your 1-pager — a concise document that captures what you're looking for and what you bring to the field. This is useful both for your own thinking and as something you can share with potential employers or collaborators.
We don't expect this to be super detailed or a perfect representation. Treat this as a first draft, and something that's _good enough_ to help you get started.
The optional resources offer additional perspective on building a career in AI safety from independent research to landing roles at frontier AI companies.
Exercise — Create your 1-pager.
Spend ~1 hour on this.
Once you're finished, 1) make the Google Doc shareable, 2) submit it here, and 3) share it with your group in Slack!
Next steps: Apply to roles
AI safety orgs are hiring for people who can hit the ground running. You don't need years of AI safety experience. A common mistake is waiting until you feel "ready". If you have relevant professional experience, start applying now.
Next steps: Technical fellowships
Fellowships are a common entry point into technical AI safety research. They offer mentorship, structure, and a chance to test your fit before committing to a full-time role. If you're unsure whether you're ready, apply anyway.
Next steps: Policy fellowships
Technical understanding is valuable in policy roles. If you're drawn to shaping how AI is governed rather than building safety techniques directly, these fellowships can help you pivot.
Next steps: Other fellowships
Not all paths fit neatly into "technical" or "policy". These fellowships serve specific niches: journalists, interdisciplinary researchers, or those exploring adjacent problems like biosecurity. Worth a look if your background is unusual.