DeepSeek-R1-distill-Qwen-7B
ToolFirst seen 9/3/2026
Last seen 9/3/2026
Evidence 5 chunks
NEIGHBORHOOD
4 nodes · 5 edgesgraph · DeepSeek-R1-distill-Qwen-7B · depth=1
RELATIONSHIPS
3 connectionsThe paper uses and evaluates DeepSeek-R1-distill-Qwen-7B as its base small language model.
DeepSeek-R1-distill-Qwen-7B is fine-tuned using SFT on verification reasoning traces.
DeepSeek-R1-distill-Qwen-7B is trained with GRPO-SMu to achieve best results.