Takeaways From Our Robust Injury Classifier Project [Redwood Research]

04/01/2025 12 min

Listen "Takeaways From Our Robust Injury Classifier Project [Redwood Research]"

Episode Synopsis

With the benefit of hindsight, we have a better sense of our takeaways from our first adversarial training project (paper). Our original aim was to use adversarial training to make a system that (as far as we could tell) never produced injurious completions. If we had accomplished that, we think it would have been the first demonstration of a deep learning system avoiding a difficult-to-formalize catastrophe with an ultra-high level of reliability. Presumably, we would have needed to invent novel robustness techniques that could have informed techniques useful for aligning TAI. With a successful system, we also could have performed ablations to get a clear sense of which building blocks were most important. Alas, we fell well short of that target. We still saw failures when just randomly sampling prompts and completions. Our adversarial training didn’t reduce the random failure rate, nor did it eliminate highly egregious failures (example below). We also don’t think we've successfully demonstrated a negative result, given that our results could be explained by suboptimal choices in our training process. Overall, we’d say this project had value as a learning experience but produced much less alignment progress than we hoped.Source:https://www.alignmentforum.org/posts/n3LAgnHg6ashQK3fF/takeaways-from-our-robust-injury-classifier-project-redwoodNarrated for AI Safety Fundamentals by TYPE III AUDIO.---A podcast by BlueDot Impact.Learn more on the AI Safety Fundamentals website.

More episodes of the podcast AI Safety Fundamentals

AI and Leviathan: Part I 29/09/2025

d/acc: One Year Later 19/09/2025

A Playbook for Securing AI Model Weights 18/09/2025

AI Emergency Preparedness: Examining the Federal Government's Ability to Detect and Respond to AI-Related National Security Threats 18/09/2025

Resilience and Adaptation to Advanced AI 18/09/2025

Introduction to AI Control 18/09/2025

The Project: Situational Awareness 18/09/2025

Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? 18/09/2025

The Intelligence Curse 18/09/2025

AI Is Reviving Fears Around Bioterrorism. What’s the Real Risk? 12/09/2025

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

Takeaways From Our Robust Injury Classifier Project [Redwood Research]

Listen "Takeaways From Our Robust Injury Classifier Project [Redwood Research]"

Episode Synopsis

More episodes of the podcast AI Safety Fundamentals

Preparing for a Hacker Threat

Bandwidth: Broadband or Narrowband?

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD