Qwen EXL3 benchmark
Qwen EXL3 benchmark is Animated Qwen3.8-Flash-Next EXL3 benchmark on one DGX Spark. Recorded results, source data, and local PNG rendering.. It is ranked #178 on the AI Radar, in Open & local models, first seen 11 days ago and shared in 4 posts (88.2k views).
What people said about Qwen EXL3 benchmark on X
Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…
— @ViC305, 1 days ago · 263 likes · see the post
Qwen3.8-Flash-Next EXL3 just got a BIG one-Spark update. 🔥 58.8 tok/s single-stream through native ExLlamaV3.🚀 157.6 tok/s aggregate through vLLM across 8 streams.🤯 FULL 262,144 context on ONE DGX Spark. 4.05 bpw EXL3 pack. A much better serving envelope. 𝗕𝗘𝗙𝗢𝗥𝗘 → 𝗡𝗢𝗪 Previous public headline: 47.6 tok/s…
— @ViC305, 4 days ago · 158 likes · see the post
Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…
— @ViC305, 11 days ago · 58 likes · see the post
Alternatives to Qwen EXL3 benchmark
- Pirate Face — The un-bannable, checksum-verified mirror for open models. Claim your handle before a squatter does.
- atomic.chat — Run AI models locally ->
- deepseek-v4.1-flash-exl3-2x-dgx-sparks — DeepSeek v4.1 Flash EXL3 2.9 bpw for 2x DGX Sparks - MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
- bespoke-nimble-9b — We’re on a journey to advance and democratize artificial intelligence through open source and open science.
- Cactus Compute — It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS,…
- orcabonsai-27b-uncensored — Open source
Qwen EXL3 benchmark in numbers
- Rank on the AI Radar: #178 of 1189
- Shared in 4 posts by 2 accounts: @ViC305, @Blackwellboy
- 88.2k views on those posts
- First seen 11 days ago, last shared 1 days ago
- Pricing seen by Jev: free
- Market: Open & local models
FAQ
What is Qwen EXL3 benchmark?
Animated Qwen3.8-Flash-Next EXL3 benchmark on one DGX Spark. Recorded results, source data, and local PNG rendering. It was first shared on X 11 days ago and is ranked #178 on the AI Radar.
Is Qwen EXL3 benchmark free?
Yes, it is free to use.
Who shared Qwen EXL3 benchmark?
2 accounts on X, including @ViC305, @Blackwellboy, in 4 posts totalling 88.2k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.1k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:00 UTC. Full method.