Group Relative Policy Optimization
TechniqueFirst seen 9/3/2026
Last seen 9/3/2026
Evidence 6 chunks
NEIGHBORHOOD
4 nodes · 4 edgesgraph · Group Relative Policy Optimization · depth=1
RELATIONSHIPS
3 connectionsGRPO-SMu modifies and improves upon conventional GRPO by diversifying input states.
GRPO is a specific RL technique applied to language model training.
DeepSeekMath: pushing the limits of mathematical reasoning in open language models ← introduces 100% 1e
The DeepSeekMath paper originally proposed the GRPO technique.