qwen38-exl3-dflash2
qwen38-exl3-dflash2 is Qwen3.8-27B EXL3 (4.00 bpw) + DFlash2 speculative decoding for ExLlamaV3, validated at 262k context on a 24 GB RTX 3090 - r0b0tlab/qwen38-exl3-dflash2. It is ranked #108 on the AI Radar, in Open & local models, first seen 3 days ago and shared in 2 posts (21.6k views).
GitHub - r0b0tlab/qwen38-exl3-dflash2: Qwen3.8-27B EXL3 (4.00 bpw) + DFlash2 speculative decoding for ExLlamaV3, validated at 262k context on a 24 GB RTX 3090 · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} r0b0tlab / qwen38-exl3-dflash2 Public Notifications You must be signed in to change notification settings Fork 2 Star 18…
What people said about qwen38-exl3-dflash2 on X
r0b0tlab/Qwen3.8-27B-4.0bpw-EXL3 + DFlash2-EXL3 + optimized ExLlamaV3 runtime! Measured on one @NVIDIAAI RTX 3090 FE ⚡ 152.4 tok/s at 8K context 🧠 262K context capacity 💾 22.7 GiB peak VRAM runtime, weights, evals: https://github.com/r0b0tlab/qwen38-exl3-dflash2
— @mr_r0b0t, 3 days ago · 240 likes · see the post
You can run these quants in buun-llama btw. They work really well, I was just polishing off getting safetensor drafters to work. buun/llama-server -hf r0b0tlab/Qwen3.8-27B-4.0bpw-EXL3 -md DFlash2-EXL3/ Feel free to use your normal favorite drafter ggufs too, you can mix it up.
— @spiritbuun, 1 days ago · 59 likes · see the post
Alternatives to qwen38-exl3-dflash2
- Pirate Face — The un-bannable, checksum-verified mirror for open models. Claim your handle before a squatter does.
- atomic.chat — Run AI models locally ->
- deepseek-v4.1-flash-exl3-2x-dgx-sparks — DeepSeek v4.1 Flash EXL3 2.9 bpw for 2x DGX Sparks - MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
- bespoke-nimble-9b — We’re on a journey to advance and democratize artificial intelligence through open source and open science.
- Cactus Compute — It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS,…
- orcabonsai-27b-uncensored — Open source
qwen38-exl3-dflash2 in numbers
- Rank on the AI Radar: #108 of 1183
- Shared in 2 posts by 2 accounts: @mr_r0b0t, @spiritbuun
- 21.6k views on those posts
- First seen 3 days ago, last shared 1 days ago
- Pricing seen by Jev: open source
- Market: Open & local models
FAQ
What is qwen38-exl3-dflash2?
Qwen3.8-27B EXL3 (4.00 bpw) + DFlash2 speculative decoding for ExLlamaV3, validated at 262k context on a 24 GB RTX 3090 - r0b0tlab/qwen38-exl3-dflash2 It was first shared on X 3 days ago and is ranked #108 on the AI Radar.
Is qwen38-exl3-dflash2 free?
It is open source.
Who shared qwen38-exl3-dflash2?
2 accounts on X, including @mr_r0b0t, @spiritbuun, in 2 posts totalling 21.6k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.1k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 17:00 UTC. Full method.