bonsai2-small-gpu
bonsai2-small-gpu is run ternary bonsai 2 27b well on the gpus people own: serve lines per vram tier, a 1.5x decode kernel for the prismml fork, the qwen 3.8 mtp head grafted back, sweeps by pr - sudoingX/bonsai2-small.... It is ranked #193 on the AI Radar, in Open & local models, first seen 2 h ago and shared in 2 posts (2.8k views).
GitHub - sudoingX/bonsai2-small-gpu: run ternary bonsai 2 27b well on the gpus people own: serve lines per vram tier, a 1.5x decode kernel for the prismml fork, the qwen 3.8 mtp head grafted back, sweeps by pr · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} sudoingX / bonsai2-small-gpu Public Notifications You must be signed…
What people said about bonsai2-small-gpu on X
you can now run qwen 3.8 27b, bonsai 2 at 40 tok/s on a single rtx 3060 or any 12gb card, full agentic loop, the whole 262k context window resident. let that sink in. i am saying this because i ran it myself and built things with it. a 77 minute agent build on that card, a working file at the end, hermes agent on top…
— @sudoingX, 2 h ago · 58 likes · see the post
the most owned gpu on steam is a five year old rtx 3060 12gb vram. today in 2026 that card runs a qwen 3.8 27b dense: > 26 tok/s on prismml's stock fork > 40 tok/s with my kernel fix, same file > 50 tok/s with the mtp head grafted back, lossless, the words do not change (releasing) > 262k context resident, 11.7 of 12…
— @sudoingX, 18 min ago · 0 likes · see the post
Alternatives to bonsai2-small-gpu
- Pirate Face — The un-bannable, checksum-verified mirror for open models. Claim your handle before a squatter does.
- Intern-S2-397B — 🤖 ModelScope
- bespoke-nimble-9b — We’re on a journey to advance and democratize artificial intelligence through open source and open science.
- atomic.chat — Run AI models locally ->
- deepseek-v4.1-flash-exl3-2x-dgx-sparks — DeepSeek v4.1 Flash EXL3 2.9 bpw for 2x DGX Sparks - MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
- orcabonsai-27b-uncensored — Open source
bonsai2-small-gpu in numbers
- Rank on the AI Radar: #193 of 1344
- Shared in 2 posts by 1 account: @sudoingX
- 2.8k views on those posts
- First seen 2 h ago, last shared 18 min ago
- Pricing seen by Jev: open source
- Market: Open & local models
FAQ
What is bonsai2-small-gpu?
run ternary bonsai 2 27b well on the gpus people own: serve lines per vram tier, a 1.5x decode kernel for the prismml fork, the qwen 3.8 mtp head grafted back, sweeps by pr - sudoingX/bonsai2-small... It was first shared on X 2 h ago and is ranked #193 on the AI Radar.
Is bonsai2-small-gpu free?
It is open source.
Who shared bonsai2-small-gpu?
1 account on X, including @sudoingX, in 2 posts totalling 2.8k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 31.7k posts from 4.8k X accounts over the last 14 days, 1.3k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 03:20 UTC. Full method.