Big update to the @Alibaba_Qwen Qwen3.8-Flash-Next single DGX Spark recipe!
This is a AI post classified by Jev as AI infra & evals (a tool drop), kept by the AI Radar because it carries real work, not commentary.
Big update to the @Alibaba_Qwen Qwen3.8-Flash-Next single DGX Spark recipe! 𝗪𝗵𝗮𝘁'𝘀 𝗻𝗲𝘄 🎁 • Measured on one DGX Spark, 262K context, MTP k=3, aggregate tok/s at 1 / 2 / 4 / 8 streams: Prose: 38.0 / 61.1 / 89.2 / 117.4 Code: 53.8 / 87.5 / 131.8 / 180.2 • Long context holds: MTP keeps working at a 185K-token prompt (35.4 tok/s decode), prefill ~2,000 tok/s from 4K to 185K • ~1M-token KV pool at the full 262K context (FP8 KV) • NVIDIA's official NVFP4 checkpoint now runs on one Spark, with chat, tool calls and vision working • 24/7 mode: a supervisor restarts the server on a crash, detects
Posted by Javier • priv/acc (4.3k followers) 1 h ago · 28 likes · 9.5k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Here's a useful use-case for @typesafeai for a change. — @iam_zachi
- This is genius. — @daniel_mac8
- Which GPU should you use for embedding workloads? — @runpod
- 🔰 The Databricks Advanced Learning Festival runs September 16 through October 14, with… — @databricks
- Q: What is a trace? — @HamelHusain
- Inference scaling part 1. — @rasbt
- AI only moves fast when your processes can keep up. See how Gearset helped BambooHR… — @Dreamforce
- cloudflare 的 AI gateway 可以直接使用 jev 这个模型,无需申请 — @madawei2699
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 16:51 UTC. Full method.