-
Efficient RL Training for LLMs with Experience Replay
Paper • 2604.08706 • Published • 24 -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Paper • 2605.30789 • Published • 26 -
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Paper • 2607.18722 • Published • 35
Chintu Kumar
chang2394
AI & ML interests
None yet
Organizations
None yet
Off policy/entropy
-
Efficient RL Training for LLMs with Experience Replay
Paper • 2604.08706 • Published • 24 -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Paper • 2605.30789 • Published • 26 -
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Paper • 2607.18722 • Published • 35
RL Advantage
models 0
None public yet
datasets 0
None public yet