GRPO (Group Relative Policy Optimization)

05/02/2025 12 min

Listen "GRPO (Group Relative Policy Optimization)"

Descargar episodio Ver en sitio original

Episode Synopsis

Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm that enhances mathematical reasoning in large language models (LLMs). It is like training students in a study group, where they learn by comparing answers without a tutor. GRPO eliminates the need for a critic model, unlike Proximal Policy Optimization (PPO), making it more resource efficient. It calculates advantages based on relative rewards within the group and directly adds KL divergence to the loss function. GRPO uses both outcome and process supervision, and can be applied iteratively, further enhancing performance. This approach is effective at improving LLMs' math skills with reduced training resources.

More episodes of the podcast Large Language Model (LLM) Talk

Kimi K2 22/07/2025

Mixture-of-Recursions (MoR) 18/07/2025

MeanFlow 10/07/2025

Mamba 10/07/2025

LLM Alignment 14/06/2025

Why We Think 20/05/2025

Deep Research 12/05/2025

vLLM 04/05/2025

Qwen3: Thinking Deeper, Acting Faster 04/05/2025

RAGEN: train and evaluate LLM agents using multi-turn RL 03/05/2025

Ver todos los episodios

ZARZA We are Zarza, the prestigious firm behind major projects in information technology.

GRPO (Group Relative Policy Optimization)

Listen "GRPO (Group Relative Policy Optimization)"

Episode Synopsis

More episodes of the podcast Large Language Model (LLM) Talk

Internet as human right and its scope

Personnel recruitment via Web

Bandwidth: Broadband or Narrowband?

Personnel recruitment via Web

Deep web or Invisible Internet

Subdomains, a glance with the experts!

Free Internet, a prediction in Nostradamus style

Educational Technology: From traditional to digital

Localhost, there’s no place like 127.0.0.1

Googling with breathtaking tricks you ignore

Gray Hat Hacking, those with ambiguous ethics…

Internet Predators on the prowl

Dot COM: The Internet’s dominant TLD