AI Radar
Support
LiveUpdated 2026-09-19 18:45 UTC

RTX 3060 users, this one is for you.

RTX 3060 users, this one is for you. @PrismML released new Ternary Bonsai 2 27B models Derived from Qwen3.8-27B, which…

This is a AI post classified by Jev as Open & local models (a model release), kept by the AI Radar because it carries real work, not commentary.

RTX 3060 users, this one is for you. @PrismML released new Ternary Bonsai 2 27B models Derived from Qwen3.8-27B, which couldn't be run reasonably on a 3060 at all. I ran their 2-bit quant PQ2_0 instead of 1-bit for this benchmark test. This is the best model to run on a 3060 🤩. llama.cpp flags: kv q8_0, fa on, np 1 max context around 88k at 11GB vram utilization Running without mtp or dflash 4k context - tg 30 t/s | peak tg 32 t/s | prefill 496 t/s 16k context - tg 29 t/s | peak tg 29 t/s | prefill 495 t/s 32k context - tg 24 t/s | peak tg 25 t/s |prefill 453 t/s 64k context - tg 18 t

Posted by AJ (7.5k followers) 1 days ago · 38 likes · 13.2k views · view the original post on X. Kept by the AI Radar as Open & local models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:45 UTC. Full method.