A 27B model in 5.95GB just ran 50 tok/s on a 12GB RTX 3060. Ternary weights plus an MTP…
This is a AI post classified by Jev as Open & local models (a model release), kept by the AI Radar because it carries real work, not commentary.
A 27B model in 5.95GB just ran 50 tok/s on a 12GB RTX 3060. Ternary weights plus an MTP graft, from @sudoingX. The arithmetic behind it: 360GB/s ÷ 5.95GB = 60 tok/s if every token reads the whole file. His fresh 50 sits at 83% of that ceiling. Ternary does not beat bandwidth, it shrinks what bandwidth has to carry. On my 273GB/s box the same file floors at 45.9 tok/s. Calculation, I have not run it yet. His 5 hour agent run at 125k context averaged 22 tok/s and wrote 2,368 lines of js. The fresh 50 is the demo number. The 22 is what an agent actually gets.
Posted by 코지베어 🐻 CozyBear (1.2k followers) 1 h ago · 0 likes · 9 views · view the original post on X. Kept by the AI Radar as Open & local models.
More AI work like this
- BIG performance update dropping today for DeepSeek v4.1 Flash running on 3x DGX Sparks 👀 — @MiaAI_lab
- This is Space Bunny. A new stealth model in @Opencode. — @superalesha
- Muse spark 1.4 has just been found in OpenCode — @MaaSonder
- Google Antigravity SDK now runs Gemma 4 26B locally on GPU via Google AI Edge LiteRT. — @0x0SojalSec
- Uncensored & Mythos-Class Agentic Qwen3.8-27B you can run locally on consumer hardware. — @0x0SojalSec
- You can now add usage credits for paid cloud models without an Ollama subscription. — @ollama
- Local Ai is about to enter the stratosphere 🤯🚀 — @ashxhart
- The Easiest way to start with Local AI Today? ODS, which is completely free (open-source… — @TheAhmadOsman
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 46k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 11:00 UTC. Full method.