Skip to content
STIMSMITH

Group Relative Policy Optimization

Technique
First seen 9/3/2026
Last seen 9/3/2026
Evidence 6 chunks

NEIGHBORHOOD

4 nodes · 4 edges
graph · Group Relative Policy Optimization · depth=1

RELATIONSHIPS

3 connections
GRPO with State Mutation ← extends 100% 2e
GRPO-SMu modifies and improves upon conventional GRPO by diversifying input states.
reinforcement learning for language models implements → 100% 2e
GRPO is a specific RL technique applied to language model training.
The DeepSeekMath paper originally proposed the GRPO technique.