AI Radar
Support
LiveUpdated 2026-09-19 17:00 UTC

exllamav3

exllamav3 — One-line default flip. Measured +7 to +13% across three prompt classes on GB10 with…

exllamav3 is One-line default flip. Measured +7 to +13% across three prompt classes on GB10 with acceptance unchanged (#2); the kernels keep fp16 arithmetic and reduction order so storage rounding is the only d.... It is ranked #84 on the AI Radar, in AI infra & evals, first seen 1 days ago and shared in 5 posts (36.9k views).

Visit github.com

What people said about exllamav3 on X

Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…

@ViC305, 1 days ago · 263 likes · see the post

Squeezing every last bit of speed out of it! Eager to try this one on the single spark, Its been chugging away on some fun H3 video gen atm but almost done. Cruz cooked here

@NeoAIForecast, 1 days ago · 23 likes · see the post

that super good speeds

@TechMDAI, 1 days ago · 9 likes · see the post

Alternatives to exllamav3

exllamav3 in numbers

FAQ

What is exllamav3?

One-line default flip. Measured +7 to +13% across three prompt classes on GB10 with acceptance unchanged (#2); the kernels keep fp16 arithmetic and reduction order so storage rounding is the only d... It was first shared on X 1 days ago and is ranked #84 on the AI Radar.

Is exllamav3 free?

It is open source.

Who shared exllamav3?

5 accounts on X, including @ViC305, @NeoAIForecast, @TechMDAI, in 5 posts totalling 36.9k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.1k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 17:00 UTC. Full method.