Listen "Reinforcement Learning Gone Wrong"
Episode Synopsis
Last week’s episode on artificial intelligence gets a huge payoff this week—we’ll explore a wonderful couple of papers about all the ways that artificial intelligence can go wrong. Malevolent actors? You bet. Collateral damage? Of course. Reward hacking? Naturally! It’s fun to think about, and the discussion starting now will have reverberations for decades to come.
https://www.technologyreview.com/s/601519/how-to-create-a-malevolent-artificial-intelligence/
http://arxiv.org/abs/1605.02817
https://arxiv.org/abs/1606.06565
https://www.technologyreview.com/s/601519/how-to-create-a-malevolent-artificial-intelligence/
http://arxiv.org/abs/1605.02817
https://arxiv.org/abs/1606.06565
More episodes of the podcast Linear Digressions
Distillation, or, How to Steal a Model
27/07/2026
Still summer break: back next week
13/07/2026
Summer break: back soon
06/07/2026