Skip to content
STIMSMITH

GRPO with State Mutation

Technique
First seen 9/3/2026
Last seen 9/3/2026
Evidence 8 chunks

NEIGHBORHOOD

8 nodes · 11 edges
graph · GRPO with State Mutation · depth=1

RELATIONSHIPS

7 connections
advantage computation uses → 100% 2e
GRPO-SMu modifies the advantage computation to handle sparse rewards.
Group Relative Policy Optimization extends → 100% 2e
GRPO-SMu modifies and improves upon conventional GRPO by diversifying input states.
tree-based branching mutation strategy uses → 100% 2e
GRPO-SMu relies on the tree-based branching mutation strategy to construct its training data.
The paper proposes GRPO-SMu as a novel RL training method.
reinforcement learning for language models implements → 100% 2e
GRPO-SMu is a specific implementation of reinforcement learning for language model training.
DeepSeek-R1-distill-Qwen-7B ← uses 100% 2e
DeepSeek-R1-distill-Qwen-7B is trained with GRPO-SMu to achieve best results.
sequential circuit verification uses → 90% 1e
GRPO-SMu shows particular improvement for sequential circuit verification.