A single RTX 3090 hit 381 tok/s on Qwen3.8-27B.
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
A single RTX 3090 hit 381 tok/s on Qwen3.8-27B. DFlash2 + lookup: 138 Optimized MTP: 114 context lookup so the model verifies 16 tokens at once when the answer already lives in the prompt. Longer verification & prefix cache: 381 tok/s The 381 number only appears when the model is reproducing or extracting from a long prompt. That’s the point. Prefix cache also drops TTFT from 22s to 0.56s on the second question against the same document. Best for RAG, document Q&A, and coding agents. Regular chat still lands around 133 tok/s. 25K-token context. https://x.com/TeksEdge/status/209059599011
Posted by Md Ismail Šojal 🕷️ (55.6k followers) 1 h ago · 14 likes · 1.2k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Live now: our Inference Engineering Track from AI Engineer World's Fair 2026. — @aiDotEngineer
- Trains your own large language model from scratch using plain PyTorch. — @tom_doerr
- Here's a useful use-case for @typesafeai for a change. — @iam_zachi
- This is genius. — @daniel_mac8
- Which GPU should you use for embedding workloads? — @runpod
- Big update to the @Alibaba_Qwen Qwen3.8-Flash-Next single DGX Spark recipe! — @jvr0x
- 🔰 The Databricks Advanced Learning Festival runs September 16 through October 14, with… — @databricks
- Q: What is a trace? — @HamelHusain
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 31.3k posts from 4.7k X accounts over the last 14 days, 1.3k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 20:46 UTC. Full method.