view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 5 days ago • 72
view article Article Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem MultiverseComputingCAI • 6 days ago • 31
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 6 days ago • 76
view article Article Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL +2 aminediroHF, qgallouedec, kashif, sergiopaniego • 17 days ago • 53
view article Article IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license ibm-research • 18 days ago • 60
ibm-granite/granite-timeseries-patchtst-fm-r2 Time Series Forecasting • 0.4B • Updated 18 days ago • 187k • 16
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 24 days ago • 73
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 24 days ago • 111
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published Aug 26 • 163
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published Aug 27 • 65
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published Aug 26 • 69