Skip to content
STIMSMITH

DeepSeek-R1-distill-Qwen-7B

Tool
First seen 9/3/2026
Last seen 9/3/2026
Evidence 5 chunks

NEIGHBORHOOD

4 nodes · 5 edges
graph · DeepSeek-R1-distill-Qwen-7B · depth=1

RELATIONSHIPS

3 connections
The paper uses and evaluates DeepSeek-R1-distill-Qwen-7B as its base small language model.
supervised fine-tuning uses → 100% 2e
DeepSeek-R1-distill-Qwen-7B is fine-tuned using SFT on verification reasoning traces.
GRPO with State Mutation uses → 100% 2e
DeepSeek-R1-distill-Qwen-7B is trained with GRPO-SMu to achieve best results.