AI Safety via Debate

04/01/2025 39 min

Listen "AI Safety via Debate"

Episode Synopsis

Abstract:To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and useful, but this approach can fail if the task is too complicated for a human to directly judge. To help address this concern, we propose training agents via self play on a zero sum debate game. Given a question or proposed action, two agents take turns making short statements up to a limit, then a human judges which of the agents gave the most true, useful information. In an analogy to complexity theory, debate with optimal play can answer any question in PSPACE given polynomial time judges (direct judging answers only NP questions). In practice, whether debate works involves empirical questions about humans and the tasks we want AIs to perform, plus theoretical questions about the meaning of AI alignment. We report results on an initial MNIST experiment where agents compete to convince a sparse classifier, boosting the classifier's accuracy from 59.4% to 88.9% given 6 pixels and from 48.2% to 85.2% given 4 pixels. Finally, we discuss theoretical and practical aspects of the debate model, focusing on potential weaknesses as the model scales up, and we propose future human and computer experiments to test these properties.Original text:https://arxiv.org/abs/1805.00899Narrated for AI Safety Fundamentals by Perrin Walker of TYPE III AUDIO.---A podcast by BlueDot Impact.Learn more on the AI Safety Fundamentals website.

More episodes of the podcast AI Safety Fundamentals

AI and Leviathan: Part I 29/09/2025

d/acc: One Year Later 19/09/2025

A Playbook for Securing AI Model Weights 18/09/2025

AI Emergency Preparedness: Examining the Federal Government's Ability to Detect and Respond to AI-Related National Security Threats 18/09/2025

Resilience and Adaptation to Advanced AI 18/09/2025

Introduction to AI Control 18/09/2025

The Project: Situational Awareness 18/09/2025

Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? 18/09/2025

The Intelligence Curse 18/09/2025

AI Is Reviving Fears Around Bioterrorism. What’s the Real Risk? 12/09/2025

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

AI Safety via Debate

Listen "AI Safety via Debate"

Episode Synopsis

More episodes of the podcast AI Safety Fundamentals

Googling with breathtaking tricks you ignore

Bandwidth: Broadband or Narrowband?

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD