mradermacher/Qwen3-8B-MiMo-Music-GRPO-GGUF Reinforcement Learning • 8B • Updated about 14 hours ago • 245