Agensh: Scaling Organizational Intelligence to 1,024 Agents Paper • 2609.26781 • Published 4 days ago • 18
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal Paper • 2609.21996 • Published 8 days ago • 9
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles Paper • 2609.22220 • Published 24 days ago • 7
Grounded Action Model: 3D Grounding as a Foundation for Robotics Paper • 2609.23863 • Published 6 days ago • 87
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 8 days ago • 34
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 9 days ago • 53