CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 13 days ago • 138
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 16 days ago • 77
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 21 days ago • 64
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 20 days ago • 65
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Paper • 2607.26769 • Published Jul 29 • 25
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification Paper • 2606.01476 • Published May 31 • 8
StableI2I: Spotting Unintended Changes in Image-to-Image Transition Paper • 2605.04453 • Published May 6 • 11
Running on CPU Upgrade Agents 1.03k Open VLM Leaderboard 🌎 1.03k VLMEvalKit Evaluation Results Collection
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Paper • 2511.20785 • Published Nov 25, 2025 • 188