Skip to content
zarza zarza
Advertisement

Q-Learning with World Models

23/08/2026 24 min

Listen "Q-Learning with World Models"

Episode Synopsis

The researchers introduce Q-Learning with World Models (QWM), a framework designed to enhance sample efficiency and performance in robotic reinforcement learning. Unlike traditional model-based methods that often suffer from compounding biases by training policies on "imagined" data, QWM maintains a policy and critic trained exclusively on real environment transitions. It leverages a learned world model specifically at test-time to conduct tree searches over potential future trajectories, allowing the agent to select actions with the highest predicted downstream value. This approach combines the predictive power of world models with the stability of grounded Q-learning to navigate complex, high-dimensional tasks. Experiments on challenging manipulation benchmarks like Robomimic and LIBERO demonstrate that QWM significantly outperforms existing model-free and model-based baselines. Ultimately, the framework scales effectively from state-based inputs to visual observations, providing a robust method for improving online reinforcement learning.

More episodes of the podcast Best AI papers explained

ZARZA Studio — Your station on air today: library, music clock, schedule, studio and reports, from the browser.

Meet ZARZA Studio
on air now stations in the catalogue 1,829,025 podcasts countries