JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments Paper • 2609.35032 • Published 5 days ago
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation Paper • 2607.10079 • Published Jul 16
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism Paper • 2607.12640 • Published Jul 14
How Output Format Confounds Data Quality and Capability in Instruction Tuning Paper • 2609.02015 • Published Sep 2
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition Paper • 2605.02094 • Published May 3
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities Paper • 2607.11197 • Published Jul 13
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations Paper • 2603.16506 • Published Mar 18 • 2
Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents Paper • 2604.17019 • Published Apr 18 • 1
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning Paper • 2601.19204 • Published Jan 27
DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors Paper • 2512.21054 • Published Dec 24, 2025
Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video Paper • 2511.11725 • Published Nov 13, 2025
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver Paper • 2604.08377 • Published Apr 9 • 226
AV-Deepfake1M Collection Datasets and Papers of 1M Deepfake Challenges • 5 items • Updated Aug 3, 2025 • 5