vllm-exl3
vllm-exl3 is Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights - vcruz305/vllm-exl3. It is ranked #1184 on the AI Radar, in AI infra & evals, first seen 12 days ago and shared in 3 posts (23.6k views).
What people said about vllm-exl3 on X
Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page…
— @ViC305, 11 days ago · 58 likes · see the post
Qwen3.8-Flash-Next hits 36.4 tok/s with built-in MTP on ONE DGX Spark. ⚡ Turboderp’s 3.05 bpw EXL3 pack, served through my vllm-exl3 plugin. The full model occupies 78.6 GiB on device, including the packed n-gram table. 𝗧𝗛𝗘 𝗡-𝗚𝗥𝗔𝗠 𝗧𝗔𝗕𝗟𝗘 𝗦𝗧𝗔𝗬𝗦 𝗣𝗔𝗖𝗞𝗘𝗗 30.4 GiB at 5 bits, with rows reconstructed…
— @ViC305, 12 days ago · 47 likes · see the post
⚡ Qwen3.8-Flash-Next: ~3-BIT EXL3 hits 47.6 tok/s median decode on ONE DGX Spark, 1.45× my 4-bit-class Q4_K_M GGUF setup. 🚀 Thinking decode is 1.60× faster. Prefill is up to 2.05× faster. MTP is ON on both sides. Same model. Same physical GB10. Same 65,536-token context. One request at a time. 𝗦𝗣𝗘𝗘𝗗: 𝗘𝗫𝗟𝟯 /…
— @ViC305, 10 days ago · 22 likes · see the post
Alternatives to vllm-exl3
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
vllm-exl3 in numbers
- Rank on the AI Radar: #1184 of 1190
- Shared in 3 posts by 1 account: @ViC305
- 23.6k views on those posts
- First seen 12 days ago, last shared 10 days ago
- Pricing seen by Jev: open source
- Market: AI infra & evals
FAQ
What is vllm-exl3?
Serve EXL3 (ExLlamaV3 trellis) quantized models on vLLM fork runtimes — any architecture, mixed per-layer bitrates, composable with source-format non-routed weights - vcruz305/vllm-exl3 It was first shared on X 12 days ago and is ranked #1184 on the AI Radar.
Is vllm-exl3 free?
It is open source.
Who shared vllm-exl3?
1 account on X, including @ViC305, in 3 posts totalling 23.6k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.