GRPO with State Mutation
TechniqueFirst seen 9/3/2026
Last seen 9/3/2026
Evidence 8 chunks
NEIGHBORHOOD
8 nodes · 11 edgesgraph · GRPO with State Mutation · depth=1
RELATIONSHIPS
7 connectionsGRPO-SMu modifies the advantage computation to handle sparse rewards.
GRPO-SMu modifies and improves upon conventional GRPO by diversifying input states.
GRPO-SMu relies on the tree-based branching mutation strategy to construct its training data.
The paper proposes GRPO-SMu as a novel RL training method.
GRPO-SMu is a specific implementation of reinforcement learning for language model training.
DeepSeek-R1-distill-Qwen-7B is trained with GRPO-SMu to achieve best results.
GRPO-SMu shows particular improvement for sequential circuit verification.