We're excited to release BananaAll, our SLM Super App.
It allows you to do EVERYTHING you need to do to trains SLMs in a single app, no terminal, no 30 chrome tabs.
The train tab allows you to train models, select datasets from presets, and use other ones with auto mapping, model size slider, it automatically generates a training script for you.
Then after you've trained the model or want to compare it to competitors, the evaluation tab, run ARC EASY, ARC Challenge, Hellaswag, PIQA, Arithmark 3, BananaMind Base Bench and more! Simple Results screen.
And lastly the inference tab, run your trained models or others.
Normally you would need seperate apps or scripts for that, but the BananaAll Super App lets you do all of that in a single app.
We also trained a small 2.5M parameter model on 200M tokens of Fineweb edu, The results: BananaMind Base Bench 854 and 53% on PIQA. On only 200M tokens.
Run it on defaults and it takes 244 s. Switch to 3 steps and it's 48.6 s. Add VAE tiling and it's 46.4 s.
The biggest culprit was the default. Z-Image Turbo is distilled to paint in few strokes, but the tool's default is 20. We were throwing away 5× for no reason. So were we, at first.
3 is the floor. Put 4 and 3 side by side and you cannot tell them apart. At 2 it collapses — water droplets and wood grain vanish, and the surface turns cloth-like.
The original was basically my own notepad that I decided to post before losing another version somewhere on my desktop lol. This one keeps that idea, with a much bigger list and a bit more organisation.
There are frameworks for building your own agents, coding tools you can actually open and use, research agents, browser automation, voice agents, and a dedicated section for agentic harnesses. Also included the useful boring stuff: testing, tracing, permissions and checking whether a project is still maintained.
This is a research directory, not a benchmark or a claim that every project has been personally tested. The main entries have public documentation and recent repository or release activity. Smaller projects and slower moving tools are labelled separately. I do my best to actually read through community feedback and real world user write ups when I research and create these lists / knowledge bases etc but as always, do your own research and go in with an open mind when you test out stuff :D
....that's what keeps it fun (for me anyway haha!)
Love always, it's a crazy world atm and stuff is moving so quickly, be kind to eachother and share knowledge and we must just get through this all good <3
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
This is the first fine tune to exceed 730 "arc-c" ("735": 144 pts higher than Qwen 3.8 27B) AND 880 ARC-E (The OpenAI, Claude and Gemini "zone of intelligence") in 8 bit and over 718 arc-c in 4 bit.
This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality.
In other words while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more.
This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants.
PS: There are 29 additional quant repos as of this writing too, as well NVFP4 and many more as well.
This is one of 10+ Qwen 3.8 27B at or above ARC-C of 717 (all 10 exceed all core benchmarks of Qwen 3.8, 3.6 and 3.5 27B and 35B-A3B versions) - you can see the complete project and some of the training here :