GLM 5.3 Flash EXL3 on 2x DGX Spark: a new cache update is merged!
This is a AI post classified by Jev as AI infra & evals (a launch), kept by the AI Radar because it carries real work, not commentary.
GLM 5.3 Flash EXL3 on 2x DGX Spark: a new cache update is merged! A full length request now claims 33-44% less of the KV cache pool, and one repeat prompt case got 90% faster!!! It's opt-in, read more below 👇
Posted by netrunner (7k followers) 1 h ago · 49 likes · 4.8k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- We want to sponsor Rust maintainers and open-source projects across web development,… — @encoredotdev
- Spent a bit of time tuning NCCL for TP=4 GLM 5.3 Flash on 4x DGX Sparks. Bumped to… — @mmastrac
- It has been a really interesting experience to build an eval suite for System One Models… — @morganlinton
- We got @DeepSeek-V4.1-Flash running in EXL3 on 4× Sparks! — @cfontes
- Qualcomm announces it most powerful Snapdragon 8 Elite Extreme Gen 6 chip built for… — @CNET
- The https://MLX.fast community has finally doubled the speed of Qwen 3.8 Flash Next to… — @TheDavidTai
- Have you ever had a local coding agent flying at the start of a session and crawling an… — @jvr0x
- 🚀 MCDMA is working between my Mac Studio & DGX Sparks, and now it's time to bring it to… — @ashxhart
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.4k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 14:57 UTC. Full method.