DAPO: an open-source LLM reinforcement learning system at scale
PaperFirst seen 9/3/2026
Last seen 9/3/2026
Evidence 2 chunks
NEIGHBORHOOD
3 nodes · 2 edgesgraph · DAPO: an open-source LLM reinforcement learning system at scale · depth=1
RELATIONSHIPS
2 connectionsThe paper cites DAPO for its token-level loss approach and adapts it for sparse reward handling.
The DAPO paper introduces the DAPO open-source RL system.