“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

15/12/2025 12 min

Listen "“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan"

Descargar episodio Ver en sitio original

Episode Synopsis

Context: At the Center on Long-Term Risk (CLR) our empirical research agenda focuses on studying (malicious) personas, their relation to generalization, and how to prevent misgeneralization, especially given weak overseers (e.g., undetected reward hacking) or underspecified training signals. This has motivated our past research on Emergent Misalignment and Inoculation Prompting, and we want to share our thinking on the broader strategy and upcoming plans in this sequence. TLDR: Ensuring that AIs behave as intended out-of-distribution is a key open challenge in AI safety and alignment. Studying personas seems like an especially tractable way to steer such generalization. Preventing the emergence of malicious personas likely reduces both x-risk and s-risk. Why was Bing Chat for a short time prone to threatening its users, being jealous of their wife, or starting fights about the date? What makes Claude Opus 3 special, even though it's not the smartest model by today's standards? And why do models sometimes turn evil when finetuned on unpopular aesthetic preferences , or when they learned to reward hack? We think that these phenomena are related to how personas are represented in LLMs, and how they shape generalization. Influencing generalization towards desired outcomes. Many technical AI safety [...] ---Outline:(01:32) Influencing generalization towards desired outcomes.(02:43) Personas as a useful abstraction for influencing generalization(03:54) Persona interventions might work where direct approaches fail(04:49) Alignment is not a binary question(05:47) Limitations(07:57) Appendix(08:01) What is a persona, really?(09:17) How Personas Drive Generalization The original text contained 4 footnotes which were omitted from this narration. ---
First published:
December 15th, 2025

Source:
https://www.lesswrong.com/posts/kCtyhHfpCcWuQkebz/a-case-for-model-persona-research
---
Narrated by TYPE III AUDIO.
---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

More episodes of the podcast LessWrong (30+ Karma)

“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg 15/12/2025

“Do you love Berkeley, or do you just love Lighthaven conferences?” by Screwtape 15/12/2025

“The Axiom of Choice is Not Controversial” by GenericModel 15/12/2025

“A high integrity/epistemics political machine?” by Raemon 14/12/2025

“No, Americans Don’t Think Foreign Aid Is 26% of the Budget” by Julius 14/12/2025

“The Inevitable Evolution of AI Agents” by Steven McCulloch 14/12/2025

“Why did I believe Oliver Sacks?” by Eye You 14/12/2025

“Conditional On Long-Range Signal, Ising Still Factors Locally” by johnswentworth, David Lorell 14/12/2025

[Linkpost] “Wages under superintelligence” by Zachary Brown 14/12/2025

“Filler tokens don’t allow sequential reasoning” by Brendan Long 14/12/2025

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

Listen "“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan"

Episode Synopsis

More episodes of the podcast LessWrong (30+ Karma)

White Hat Hacking, Ethical Hackers…

CAPTCHA for human verification!

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD