AI Radar
Support
LiveUpdated 2026-09-19 16:24 UTC

qwen3.8-flash-next-exl3-dgx-spark-recipe

qwen3.8-flash-next-exl3-dgx-spark-recipe — Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM +…

qwen3.8-flash-next-exl3-dgx-spark-recipe is Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe. It is ranked #47 on the AI Radar, in AI infra & evals, first seen 11 days ago and shared in 11 posts (105.7k views).

Visit github.com

What people said about qwen3.8-flash-next-exl3-dgx-spark-recipe on X

Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…

@ViC305, 1 days ago · 263 likes · see the post

Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s…

@ViC305, 4 days ago · 158 likes · see the post

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…

@ViC305, 11 days ago · 58 likes · see the post

Alternatives to qwen3.8-flash-next-exl3-dgx-spark-recipe

qwen3.8-flash-next-exl3-dgx-spark-recipe in numbers

FAQ

What is qwen3.8-flash-next-exl3-dgx-spark-recipe?

Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe It was first shared on X 11 days ago and is ranked #47 on the AI Radar.

Is qwen3.8-flash-next-exl3-dgx-spark-recipe free?

It is open source.

Who shared qwen3.8-flash-next-exl3-dgx-spark-recipe?

6 accounts on X, including @ViC305, @Blackwellboy, @NeoAIForecast, in 11 posts totalling 105.7k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 16:24 UTC. Full method.