Spent a bit of time tuning NCCL for TP=4 GLM 5.3 Flash on 4x DGX Sparks. Bumped to…
This is a AI post classified by Jev as AI infra & evals (a tutorial), kept by the AI Radar because it carries real work, not commentary.
Spent a bit of time tuning NCCL for TP=4 GLM 5.3 Flash on 4x DGX Sparks. Bumped to latest vLLM nightly, enabled _both_ ConnectX ports at 200gb. Might have one of the fastest DGX Spark deploys out there too. ~60 tok/s for prose, ~88 tok/s for code and >3000 tok/s for prefill.
Posted by Matt Mastracci (8.7k followers) 1 h ago · 0 likes · 211 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Your model weights are encrypted at rest. Cool. — @CoreWeave
- Not bad — @zephyr_z9
- We want to sponsor Rust maintainers and open-source projects across web development,… — @encoredotdev
- GLM 5.3 Flash EXL3 on 2x DGX Spark: a new cache update is merged! — @plotarmordev
- It has been a really interesting experience to build an eval suite for System One Models… — @morganlinton
- We got @DeepSeek-V4.1-Flash running in EXL3 on 4× Sparks! — @cfontes
- Qualcomm announces it most powerful Snapdragon 8 Elite Extreme Gen 6 chip built for… — @CNET
- The https://MLX.fast community has finally doubled the speed of Qwen 3.8 Flash Next to… — @TheDavidTai
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.5k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 15:40 UTC. Full method.