Just launched my usual Qwen 3.8 27B configs and, for a brief moment, couldn't understand…

This is a AI post classified by Jev as Open & local models (a tool drop), kept by the AI Radar because it carries real work, not commentary.
Just launched my usual Qwen 3.8 27B configs and, for a brief moment, couldn't understand why I had so little VRAM in use, like multiple gigabytes less than usual. This was, of course, due to the latest change in BeeLlama v0.4.7 Preview, where independent draft batch sizing drastically reduces VRAM wasted on buffers, with no performance impact. Here's a breakdown of the savings for MTP and DFlash 2 on 2×3090, using Qwen 3.8 27B UD-Q6_K_M with a 256k KVarN6 KV cache.
Posted by Ivan Neustroev (273 followers) 1 h ago · 1 likes · 24 views · view the original post on X. Kept by the AI Radar as Open & local models.
More AI work like this
- We’re open-sourcing the Ming-Image-0.1-Design family: — @AntLingAGI
- Introducing https://theopenfrontier.com! — @nutlope
- Meet Lanni-ni/hard_3gram_4_6_384_babylm_10m_seed44: a text-generation transformer… — @HuggingModels
- Meet Lanni-ni/hard_3gram_babylm_100m_2layer: a 100M param text generation model with a 2… — @HuggingModels
- i tired MiMo-v2.6-Flash once more to build a proper 3d robotic hummingbird prototype… — @TimJayas
- OmniEdu: open 4B/9B/27B models for K-12 learning and teaching — @HuggingPapers
- 2,304 of you pulled this model in under two days. rtx 3060, 3070, 4060, 4070 owners, you… — @sudoingX
- Jev launched just 7 days ago. — @minchoi
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 42.6k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 18:12 UTC. Full method.