Skip to content
STIMSMITH

DeepSeekMath: pushing the limits of mathematical reasoning in open language models

Paper
First seen 9/3/2026
Last seen 9/3/2026
Evidence 1 chunks

NEIGHBORHOOD

3 nodes · 2 edges
graph · DeepSeekMath: pushing the limits of mathematical reasoning in open language models · depth=1

RELATIONSHIPS

2 connections
The paper cites DeepSeekMath as the origin of the GRPO technique.
Group Relative Policy Optimization introduces → 100% 1e
The DeepSeekMath paper originally proposed the GRPO technique.