RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 5 days ago • 205
What Makes Recurrence Effective in Looped Language Models? Paper • 2609.36636 • Published 7 days ago • 10
PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond? Paper • 2609.34314 • Published 8 days ago • 4
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation Paper • 2609.37687 • Published 7 days ago • 6
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 10 days ago • 322
orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF Text Generation • 27B • Updated 4 days ago • 16.8k • 391