6 GiB. That is the entire KV bill for a 262K context window on Qwen3.8-Flash-Next,…
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
6 GiB. That is the entire KV bill for a 262K context window on Qwen3.8-Flash-Next, config math, not a benchmark. 48 layers, but only 12 run full attention, one in four. Those layers carry 2 KV heads at head_dim 256, which lands at 24 KiB per token, so 262,144 tokens is 6 GiB, 3 at 4-bit KV. The other 36 layers are linear attention with a fixed 1.5 MiB state each. 54 MiB whether the window is 2K or 262K. I keep seeing KV estimates for this thing scaled off dense models. The layer mix is the whole story.
Posted by 코지베어 🐻 CozyBear (1.2k followers) 2 h ago · 0 likes · 8 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- we've been exploring how to use Jev to build a better harness — @hwchase17
- Step-5-preview inference speed has improved, time to run more evals. — @ItsmeAjayKV
- 95 tok/s 🤯 — @MiaAI_lab
- I BUILT A ROUTING DESK FOR 188 AI MODELS TO SEE HOW MUCH OF AN 11-TOOL, $277/MO STACK I… — @gippp69
- I got tired of treating 20 ChatGPT/Copilot subscriptions like 20 separate accounts. So I… — @GeoffreyHuntley
- Can you use Jev for Evals? Yes! Remember that a LLM Judge is also classifier*. — @HamelHusain
- I built ds-router for Hermes Agent. — @btsouth
- Happy weekend! Let's forget all the troubles in daily life; just go outside to touch… — @alibaba_cloud
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.5k posts from 4.9k X accounts over the last 14 days, 1.5k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 14:15 UTC. Full method.