Skip to content
STIMSMITH

GRPO-SMu training dataset

CodeArtifact
First seen 9/3/2026
Last seen 9/3/2026
Evidence 2 chunks

NEIGHBORHOOD

4 nodes · 5 edges
graph · GRPO-SMu training dataset · depth=1

RELATIONSHIPS

3 connections
The paper uses the GRPO-SMu training dataset constructed via tree-based mutation strategy.
RL fine-tuning uses the GRPO-SMu training dataset constructed via tree-based mutations.
ScaleRTL dataset derived from → 100% 1e
The GRPO-SMu training dataset is constructed from golden RTL codes sourced from the ScaleRTL dataset.