AI Radar
Support
LiveUpdated 2026-09-19 18:08 UTC

Qwen3.8-Flash-Next Recipe for 2 @NVIDIAAI DGX Sparks just got an update!

Qwen3.8-Flash-Next Recipe for 2 @NVIDIAAI DGX Sparks just got an update! The wins 👇 ・+8.4% faster decode, no quality…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.

Qwen3.8-Flash-Next Recipe for 2 @NVIDIAAI DGX Sparks just got an update! The wins 👇 ・+8.4% faster decode, no quality change ・50.5 → 56.5 tok/s single-stream prose (71.3 code) ・Shrank the drafter's vocab 248k → 47k tokens ・8 streams: 227.7 tok/s aggregate prose, 328.6 code ・Each user still gets ~28 tok/s at 8 streams ・4x the users costs only 38% of per-stream speed ・Gains held at every concurrency: +3.8% to +12.9% ・One invisible newline stopped the server booting ・Freed 70 GB of stale RAM cache per launch, no root ・4 PRs merged, 8 issues closed, backlog 16 → 8 Thanks to all community members

Posted by Javier • priv/acc (4.3k followers) 3 days ago · 63 likes · 26.5k views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.