qwen3.8-flash-next-exl3-dgx-spark-recipe
qwen3.8-flash-next-exl3-dgx-spark-recipe is Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe. It is ranked #47 on the AI Radar, in AI infra & evals, first seen 11 days ago and shared in 11 posts (105.7k views).
What people said about qwen3.8-flash-next-exl3-dgx-spark-recipe on X
Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…
— @ViC305, 1 days ago · 263 likes · see the post
Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s…
— @ViC305, 4 days ago · 158 likes · see the post
Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…
— @ViC305, 11 days ago · 58 likes · see the post
Alternatives to qwen3.8-flash-next-exl3-dgx-spark-recipe
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
qwen3.8-flash-next-exl3-dgx-spark-recipe in numbers
- Rank on the AI Radar: #47 of 1181
- Shared in 11 posts by 6 accounts: @ViC305, @Blackwellboy, @NeoAIForecast, @TechMDAI
- 105.7k views on those posts
- First seen 11 days ago, last shared 23 h ago
- Pricing seen by Jev: open source
- Market: AI infra & evals
FAQ
What is qwen3.8-flash-next-exl3-dgx-spark-recipe?
Serve Qwen3.8-Flash-Next (turboderp ExLlamaV3 pack) on one NVIDIA DGX Spark with vLLM + vllm-exl3 - vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe It was first shared on X 11 days ago and is ranked #47 on the AI Radar.
Is qwen3.8-flash-next-exl3-dgx-spark-recipe free?
It is open source.
Who shared qwen3.8-flash-next-exl3-dgx-spark-recipe?
6 accounts on X, including @ViC305, @Blackwellboy, @NeoAIForecast, in 11 posts totalling 105.7k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 16:24 UTC. Full method.