AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Paper • 2609.40097 • Published 4 days ago • 24
Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration Paper • 2609.35342 • Published 6 days ago • 7
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 7 days ago • 550
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale Paper • 2609.38822 • Published 4 days ago • 6
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 11 days ago • 56
PUBG Ally: A Conversational Embodied Agent as an AI Teammate Paper • 2609.29837 • Published 10 days ago • 25
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 10 days ago • 28
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 10 days ago • 102
Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem Paper • 2609.30216 • Published 10 days ago • 15
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL Paper • 2609.29050 • Published 10 days ago • 13
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 19 days ago • 45
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 17 days ago • 110
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published about 1 month ago • 119