The https://MLX.fast community has finally doubled the speed of Qwen 3.8 Flash Next to…
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
The https://MLX.fast community has finally doubled the speed of Qwen 3.8 Flash Next to 82.7 tps (serial decode)! Congrats to https://www.yukon.org/solver/ercumentyildirim for being the one to finally push us over with a Split K QMV kernel!
Posted by TheDavidTai (1.1k followers) 1 h ago · 13 likes · 546 views · view the original post on X. Kept by the AI Radar as AI infra & evals. Tools mentioned: mlx.fast.
More AI work like this
- Qualcomm announces it most powerful Snapdragon 8 Elite Extreme Gen 6 chip built for… — @CNET
- Have you ever had a local coding agent flying at the start of a session and crawling an… — @jvr0x
- 🚀 MCDMA is working between my Mac Studio & DGX Sparks, and now it's time to bring it to… — @ashxhart
- 🚀 MCDMA has RDMA working between my Mac Studio and DGX Sparks. — @ashxhart
- Deploy DiffusionGemma-Jev on Cloud Run with just a single command. — @Saboo_Shubham_
- Update for GLM 5.3 Flash EXL3 (2x Spark) releasing soon! — @plotarmordev
- before fraud protection I was spending $15/hour on http://classifier.dev — @michael_chomsky
- 762 commits. 315 contributors. 104 first-timers. vLLM v0.30.0 is live. 🎉 — @vllm_project
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.3k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 13:36 UTC. Full method.