yandex/AliceAI-Foundation-80B-A3B-Base Text Generation • 81B • Updated about 4 hours ago • 5.97k • 384
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Text Generation • 33B • Updated Feb 24, 2025 • 417k • • 1.63k
Focused Transformer: Contrastive Training for Context Scaling Paper • 2307.03170 • Published Jul 6, 2023 • 12