[short] Controlled Decoding from Language Models

27/10/2023 1 min

Listen "[short] Controlled Decoding from Language Models"

Episode Synopsis

Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach for aligning language models.

https://arxiv.org/abs//2310.17022

YouTube: https://www.youtube.com/@ArxivPapers

TikTok: https://www.tiktok.com/@arxiv_papers

Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016

Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

More episodes of the podcast Arxiv Papers