view post Post 142 I put together a Space to showcase tokenization-in-the-browser libs we maintain at HFLet me know what you think! sbrandeis/tokenizers-wasm-demo See translation 🔥 1 1 + Reply
RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs Paper • 2609.12552 • Published 17 days ago • 5
view post Post 637 We’ve been working with LAION on a voice acting arena. You listen to two models doing the same scene and compare how well they pull it off.It’s ready to try now - would love to hear what you think 🙂 TTS-AGI/voice-acting-arena See translation Reply
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published Aug 20 • 12
PixRestore: Unified Image Restoration via Pixel Diffusion Transformer Paper • 2608.16793 • Published Aug 17 • 4
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding Paper • 2608.17402 • Published Aug 18 • 20
A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples Paper • 2607.29122 • Published Jul 31 • 6
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Paper • 2607.23909 • Published Jul 27 • 8
view post Post 5640 who's working on an NVFP4 version of Kimi-K3? See translation 4 replies · 🤗 11 11 👍 6 6 🚀 4 4 🔥 4 4 😔 1 1 + Reply
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach Paper • 2607.14702 • Published Jul 16
Team LEYA in 10th ABAW Competition: Multimodal Ambivalence/Hesitancy Recognition Approach Paper • 2603.12848 • Published Mar 13
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 78
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published Jul 13 • 43
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 33
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity Paper • 2607.07386 • Published Jul 8 • 12
Rank-Then-Act: Reward-Free Control from Frame-Order Progress Paper • 2607.01897 • Published Jul 2 • 7
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better Paper • 2607.04884 • Published Jul 6 • 10