Post
22
I tested every uncensored Ternary Bonsai 2 27B on the Hub, all 11 builds from 6 uploaders, including my own. Same prompts, same judge, same GPUs, every file pinned by sha256.
• Best answers: @Hikari07jp and @dealignai (~0.94 answer quality with thinking off)
• Only edit with no measurable MMLU cost: mine (±0.15 pp). Every other edit loses 0.64–2.68 pp, all p < 0.001
• Heretic and Blackfrost still refuse 11–12% of harmful prompts
• Thinking mode at 4,096 tokens: 5–31% of harmful prompts get no answer. PrismML recommends 16,384+, and I'm rerunning at that budget
Full report, charts and model-card checks: BoldingBuilds/bonsai-2-uncensored-shootout
• Best answers: @Hikari07jp and @dealignai (~0.94 answer quality with thinking off)
• Only edit with no measurable MMLU cost: mine (±0.15 pp). Every other edit loses 0.64–2.68 pp, all p < 0.001
• Heretic and Blackfrost still refuse 11–12% of harmful prompts
• Thinking mode at 4,096 tokens: 5–31% of harmful prompts get no answer. PrismML recommends 16,384+, and I'm rerunning at that budget
Full report, charts and model-card checks: BoldingBuilds/bonsai-2-uncensored-shootout