LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Paper • 2506.05260 • Published Jun 5, 2025
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Paper • 2607.23265 • Published Aug 3 • 1
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding Paper • 2607.23951 • Published Jul 27 • 1
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 11 days ago • 12
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 11 days ago • 56
EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory Paper • 2609.37923 • Published 11 days ago • 9
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 11 days ago • 115
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 12 days ago • 413
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 11 days ago • 12
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 11 days ago • 12
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Paper • 2609.06245 • Published Sep 5 • 28
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published Aug 21 • 60
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 65