deepseek-v4.1-flash-vllm-dgx-spark
deepseek-v4.1-flash-vllm-dgx-spark is DeepSeek-V4.1-Flash (552B MoE, MXFP4 experts, 1M ctx) on four NVIDIA DGX Sparks with vLLM TP4: Engram-on-disk patch, sm121 kernel build, launchers, measured numbers - tonyd2wild/DeepSeek-V4.1-Flash.... It is ranked #29 on the AI Radar, in AI infra & evals, first seen 9 days ago and shared in 20 posts (93k views).
What people said about deepseek-v4.1-flash-vllm-dgx-spark on X
UPDATE ! 🔥 DeepSeek-V4.1-Flash · 552B MoE (769B w/ Engram) · MXFP4/MXFP8 · 4x DGX Spark TP4 · vLLM + DSpark Day-0 bring-up, now on boot 10 🚀 ⚡️ 73.8 tok/s CODE, one stream (+41% vs last boot) 🔢 92 tok/s counting · 55 tables · 52 JSON · 51 math 📈 225 tok/s code at 6 streams · 260 peak aggregate · 132 avg across 8…
— @Tech2Wild, 8 days ago · 151 likes · see the post
Did A BIG Speed Run Update Last Night ! 🔓⚡️ DeepSeek V4.1-Flash on 3-4 DGX Sparks just went uncensored AND got faster thanks to @u1tra_instinct 🚀 ~190 tok/s throughput (6 streams, +13%) 💻 ~85 tok/s on code, single stream 📥 ~2,000 tok/s cold prefill (+38%) ⏱️ 0.22s to first token (-20%) 📚 1M-token context, proven…
— @Tech2Wild, 5 days ago · 96 likes · see the post
🔥 DeepSeek-V4.1-Flash 552B MoE (769B w/ Engram) · MXFP4 · 4x DGX Spark TP4 · vLLM + DSpark (Checkpoint) Benched warmed + streaming, fixed 8-category prompt set, C1 to C6: ⚡️ 40.5 tok/s single-stream on counting · 39 structured tables · 33 code · 31 math · 15 prose 🎯 DSpark k=5 runs hot: up to 5.6 of 6 draft tokens…
— @Tech2Wild, 9 days ago · 91 likes · see the post
Alternatives to deepseek-v4.1-flash-vllm-dgx-spark
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
deepseek-v4.1-flash-vllm-dgx-spark in numbers
- Rank on the AI Radar: #29 of 1181
- Shared in 20 posts by 4 accounts: @Tech2Wild, @u1tra_instinct, @Blackwellboy, @WescheNex1q
- 93k views on those posts
- First seen 9 days ago, last shared 2 h ago
- Pricing seen by Jev: open source
- Market: AI infra & evals
FAQ
What is deepseek-v4.1-flash-vllm-dgx-spark?
DeepSeek-V4.1-Flash (552B MoE, MXFP4 experts, 1M ctx) on four NVIDIA DGX Sparks with vLLM TP4: Engram-on-disk patch, sm121 kernel build, launchers, measured numbers - tonyd2wild/DeepSeek-V4.1-Flash... It was first shared on X 9 days ago and is ranked #29 on the AI Radar.
Is deepseek-v4.1-flash-vllm-dgx-spark free?
It is open source.
Who shared deepseek-v4.1-flash-vllm-dgx-spark?
4 accounts on X, including @Tech2Wild, @u1tra_instinct, @Blackwellboy, in 20 posts totalling 93k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 16:11 UTC. Full method.