AI Radar
Support
LiveUpdated 2026-09-19 18:08 UTC

My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the…

My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the ceiling. I was wrong. @mr_r0b0t…

This is a AI post classified by Jev as Open & local models (a model release), kept by the AI Radar because it carries real work, not commentary.

My RTX 4090 just hit 181 tok/s on Qwen 3.8 27B model. O Monday I thought 140 was the ceiling. I was wrong. @mr_r0b0t built a different engine. I wanted to push the card again to see if I can squeeze more out of it. The headline number moved 30%, the like-for-like number moved 4%, and the gap between them is the actual picture. Here is my full breakdown of whether the recipe has legs. A DIFFERENT ENGINE Two pieces. A better quant and a better drafter. The quant is EXL3 at 4.0 bits per weight, 15.4 GiB for the whole 27B model. The drafter is DFlash2, a block-diffusion model that proposes u

Posted by Yume_X (2.1k followers) 2 days ago · 80 likes · 5.9k views · view the original post on X. Kept by the AI Radar as Open & local models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.