Running on CPU Upgrade 87 MiMo RL Environment Explorer 🧭 87 Explore the MiMo-V2.6 RL environments and run rollouts
Running 259 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 259 Building and scaling RL environments for LLM training
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF Image-Text-to-Text • 27B • Updated 14 days ago • 2M • 1.59k
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 59
view article Article How we OCR'ed 30,000 papers using Codex, open OCR models and Jobs nielsr • Apr 7 • 62