← All resources

Week 2 — Alignment Challenges

Why is it so hard to get an AI system to do what we actually mean? We’ll look at problems like reward hacking, goal misgeneralization, and deceptive behavior, along with real examples that show why good performance during training is not the same as being safe in the real world.

Core readings

Recommended readings

Optional readings