AI Radar
Support
LiveUpdated 2026-09-23 17:53 UTC

DeepSeek-V4.1-Flash EXL3 is finally alive across 4× DGX Sparks. 🔥

DeepSeek-V4.1-Flash EXL3 is finally alive across 4× DGX Sparks. 🔥 I built the quality-first 4.75 bpw SAGE quant.…

This is a AI post classified by Jev as Open & local models (a model release), kept by the AI Radar because it carries real work, not commentary.

DeepSeek-V4.1-Flash EXL3 is finally alive across 4× DGX Sparks. 🔥 I built the quality-first 4.75 bpw SAGE quant. @cfontes and @Blackwellboy spent more than a week, 14 PRs, dozens of commits and 40+ bring-up iterations making it actually run. I don’t own one Spark, let alone four. This is what open-source collaboration looks like. 𝗧𝗛𝗘 𝗥𝗘𝗖𝗘𝗜𝗣𝗧𝗦 Single-stream decode: 30.2 tok/s prose 33.4 tok/s code 33.37 tok/s best warm sustained cell Aggregate at c=1 / 2 / 4 / 8: 27.8 / 43.9 / 57.7 / 61.5 tok/s KV budget: 3.97M tokens Context: full 1,048,576 tokens Every figure is a sustained 30-secon

Posted by Cruz (2.1k followers) 56 min ago · 0 likes · 15 views · view the original post on X. Kept by the AI Radar as Open & local models. Tools mentioned: dsv4.1-flash-exl3-4.75bpw, deepseek-v4.1-flash-exl3-dgx-spark-recipe, vllm-exl3.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.7k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 17:53 UTC. Full method.