Week 3 — Technical Safety Approaches
This week is about the tools researchers are using to make AI systems easier to understand and control. We’ll compare approaches such as interpretability, evaluations and red teaming, scalable oversight, and control—and ask what each one can realistically catch.
Core readings
- A Brief Explanation of AI Control (Scher, Oct 2024) ↗
- Can we scale human feedback for complex AI tasks? An intro to scalable oversight. (BlueDot, Mar 2024) ↗
- Tracing the thoughts of a large language model (Anthropic, Mar 2025) ↗
- Difficulties with Evaluating a Deception Detector for AIs (GDM, Dec 2025) ↗
Recommended readings
Control
- The case for ensuring that powerful AIs are controlled (Redwood Research, 2024) ↗
- AI Control: Improving Safety Despite Intentional Subversion (Redwood Research, 2024) ↗
- The AI Control LessWrong sequence ↗
- Ctrl-Z: Controlling AI Agents via Resampling (Redwood Research, 2025) ↗
- An Overview of Control Measures (Redwood, 2025) ↗
- Simple probes can catch Sleeper Agents (Anthropic, 2024) ↗
Scalable Oversight
- On scalable oversight with weak LLMs judging strong LLMs (GDM, 2024) ↗
- Weak-to-strong generalization (OpenAI, 2023) ↗
- Debating with More Persuasive LLMs Leads to More Truthful Answers (Khan et al., 2024) ↗
- Measuring Progress on Scalable Oversight for Large Language Models (Anthropic, 2022) ↗
- Debate update: obfuscated arguments problem (OpenAI, 2020) ↗
Mechanistic Interpretability
Optional readings
Control
- How to prevent collusion when using untrusted models to monitor each other (Redwood, 2024) ↗
- A Sketch of an AI Control Safety Case (UK AISI / Redwood, 2025) ↗
- Four Places Where You Can Put LLM Monitoring (Redwood, 2025) ↗
- Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs (LASR Labs, 2024) ↗
- Rob Miles YouTube explainer of the AI Control paper (Robert Miles, 2024) ↗
- Preventing Language Models From Hiding Their Reasoning (Roger & Greenblatt, Redwood Research, 2023) ↗
Scalable Oversight
Mechanistic Interpretability
- Detecting Strategic Deception Using Linear Probes (Apollo Research, 2025) ↗
- Toy Models of Superposition (Elhage et al., 2022) ↗
- On the Biology of a Large Language Model (Lindsey et al., Anthropic, 2025) ↗
- Refusal in Language Models Is Mediated by a Single Direction (Arditi et al., 2024) ↗
- Auditing Language Models for Hidden Objectives — full paper (Marks et al., Anthropic, 2025) ↗
Evals
- Detecting Misbehavior in Frontier Reasoning Models (OpenAI, 2025) ↗
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (Korbak et al., 2025) ↗
- Reasoning Models Don't Always Say What They Think (Chen et al., Anthropic/OpenAI, 2025) ↗
- EvilGenie: A Reward Hacking Benchmark (Gabor et al., 2025) ↗
- Sabotage Evaluations for Frontier Models (Benton et al., Anthropic, 2024) ↗
- AI Sandbagging (van der Weij et al., 2024) ↗