762 commits. 315 contributors. 104 first-timers. vLLM v0.30.0 is live. 🎉
This is a AI post classified by Jev as AI infra & evals (a launch), kept by the AI Radar because it carries real work, not commentary.
762 commits. 315 contributors. 104 first-timers. vLLM v0.30.0 is live. 🎉 Highlights: 🤖 Hybrid-attention hot paths: Kimi K3 streamlines KDA, AttnRes, and MLA; DeepSeek-V4.1-Flash adds MXFP8 KV and async Engram; Qwen3.8-Flash-Next fuses QSA/PLE and cuts sparse-GQA overhead 🗄️ HiSparse adds a host tier beneath sparse-MLA decode; under GPU pressure, only top-k misses return to a per-request hot buffer 🛠️ Model Runner V2 brings EAGLE3-style drafts to pipeline parallelism and extends adaptive verification to every draft-model speculator through online acceptance estimation (#50514, #52228) 🖋️ Dual
Posted by vLLM (49.9k followers) 1 h ago · 48 likes · 2.1k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Update for GLM 5.3 Flash EXL3 (2x Spark) releasing soon! — @plotarmordev
- before fraud protection I was spending $15/hour on http://classifier.dev — @michael_chomsky
- Building my Jev benchmark at VulcanBench, and I don't think I have the eval suite quite… — @morganlinton
- GPT-6 Luna is almost free at API prices. — @StatsWire
- - Lightmatter officially introduced Passage L20 CPX on September 17 and joined the Open… — @PhotonCap
- how do you upgrade an endpoint to a new model while it is serving millions of prod… — @zainhas
- WHAAAAT?! There was this much slack left?! — @teortaxesTex
- New on Dedicated Model Inference: canary rollouts. — @togethercompute
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 43.8k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 04:27 UTC. Full method.