Tree-based Group Policy Optimization for LLM Agents

26/09/2025 15 min

Listen "Tree-based Group Policy Optimization for LLM Agents"

Descargar episodio Ver en sitio original

Episode Synopsis

The September 25 2025 paper introduces **Tree-based Group Relative Policy Optimization (Tree-GRPO)**, a new reinforcement learning (RL) method designed to enhance the agentic capabilities of large language models (LLMs) in multi-turn tasks where supervision is typically sparse. Tree-GRPO addresses the challenges of sparse rewards and heavy rollout costs associated with existing chain-based RL by employing a **tree-search sampling strategy** where each node represents a complete agent interaction step, allowing for prefix sharing and reduced budget use. This tree structure inherently creates **finer-grained process supervision signals** from outcome rewards, a mechanism shown to be structurally equivalent to step-level direct preference learning. Empirical results across multiple datasets demonstrate that the tree-based approach **consistently achieves higher performance with less rollout budget** compared to chain-based methods.Source:https://arxiv.org/pdf/2509.21240

More episodes of the podcast AI: post transformers

Attention with a bias 17/01/2026

Squisher: Approximating the Fisher Information Matrix and use cases 17/01/2026

NVIDIA: TTT-E2E: Unlocking Long-Context Learning via End-to-End Test-Time Training 17/01/2026

Scaling laws: long context length and in context learning 17/01/2026

DeepSeek Engram: Scaling Large Language Models via Conditional Memory Lookup 14/01/2026

PageANN: Scalable Disk ANNS with Page-Aligned Graphs 07/12/2025

NeurIPS 2025: Homogeneous Keys, Heterogeneous Values 04/12/2025

NeurIPS 2025: Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free 29/11/2025

NeurIPS 2025: Large Language Diffusion Models 29/11/2025

NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example 29/11/2025

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

Tree-based Group Policy Optimization for LLM Agents

Listen "Tree-based Group Policy Optimization for LLM Agents"

Episode Synopsis

More episodes of the podcast AI: post transformers

Positive Attitude, Share your ZARZA Attitude!

Dot COM: The Internet’s dominant TLD

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD