OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 15 days ago • 150
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 23 days ago • 707
Running on Zero MCP 3.19k Wan2.2 14B Preview 🐌 3.19k generate a video from an image with a text prompt
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 78