Skip to content
STIMSMITH

DAPO: an open-source LLM reinforcement learning system at scale

Paper
First seen 9/3/2026
Last seen 9/3/2026
Evidence 2 chunks

NEIGHBORHOOD

3 nodes · 2 edges
graph · DAPO: an open-source LLM reinforcement learning system at scale · depth=1

RELATIONSHIPS

2 connections
The paper cites DAPO for its token-level loss approach and adapts it for sparse reward handling.
DAPO introduces → 100% 2e
The DAPO paper introduces the DAPO open-source RL system.