exllamav3
exllamav3 is One-line default flip. Measured +7 to +13% across three prompt classes on GB10 with acceptance unchanged (#2); the kernels keep fp16 arithmetic and reduction order so storage rounding is the only d.... It is ranked #84 on the AI Radar, in AI infra & evals, first seen 1 days ago and shared in 5 posts (36.9k views).
What people said about exllamav3 on X
Qwen3.8-Flash-Next -> 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲…
— @ViC305, 1 days ago · 263 likes · see the post
Squeezing every last bit of speed out of it! Eager to try this one on the single spark, Its been chugging away on some fun H3 video gen atm but almost done. Cruz cooked here
— @NeoAIForecast, 1 days ago · 23 likes · see the post
that super good speeds
— @TechMDAI, 1 days ago · 9 likes · see the post
Alternatives to exllamav3
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
exllamav3 in numbers
- Rank on the AI Radar: #84 of 1183
- Shared in 5 posts by 5 accounts: @ViC305, @NeoAIForecast, @TechMDAI, @Blackfrost_AI
- 36.9k views on those posts
- First seen 1 days ago, last shared 1 days ago
- Pricing seen by Jev: open source
- Market: AI infra & evals
FAQ
What is exllamav3?
One-line default flip. Measured +7 to +13% across three prompt classes on GB10 with acceptance unchanged (#2); the kernels keep fp16 arithmetic and reduction order so storage rounding is the only d... It was first shared on X 1 days ago and is ranked #84 on the AI Radar.
Is exllamav3 free?
It is open source.
Who shared exllamav3?
5 accounts on X, including @ViC305, @NeoAIForecast, @TechMDAI, in 5 posts totalling 36.9k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.1k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 17:00 UTC. Full method.