“On Deliberative Alignment” by Zvi

11/02/2025 12 min

Listen "“On Deliberative Alignment” by Zvi"

Descargar episodio Ver en sitio original

Episode Synopsis

Not too long ago, OpenAI presented a paper on their new strategy of Deliberative Alignment.
The way this works is that they tell the model what its policies are and then have the model think about whether it should comply with a request.
This is an important transition, so this post will go over my perspective on the new strategy.
Note the similarities, and also differences, with Anthropic's Constitutional AI.

How Deliberative Alignment Works

We introduce deliberative alignment, a training paradigm that directly teaches reasoning LLMs the text of human-written and interpretable safety specifications, and trains them to reason explicitly about these specifications before answering.

We used deliberative alignment to align OpenAI's o-series models, enabling them to use chain-of-thought (CoT) reasoning to reflect on user prompts, identify relevant text from OpenAI's internal policies, and draft safer responses.
Our approach achieves highly precise [...] ---Outline:(00:29) How Deliberative Alignment Works(03:27) Why This Worries Me(07:49) For Mundane Safety It Works Well ---
First published:
February 11th, 2025

Source:
https://www.lesswrong.com/posts/CJ4yywLBkdRALc4sT/on-deliberative-alignment
---
Narrated by TYPE III AUDIO.
---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

“On Deliberative Alignment” by Zvi

Listen "“On Deliberative Alignment” by Zvi"

Episode Synopsis

More episodes of the podcast LessWrong posts by zvi

CAPTCHA for human verification!

Preparing for a Hacker Threat

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD