glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark
glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark is GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full.... It is ranked #67 on the AI Radar, in AI infra & evals, first seen 1 h ago and shared in 6 posts (4.8k views).
GitHub - tonyd2wild/GLM-5.3-Flash-NVFP4-1M-KV-4x-DGX-Spark: GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full patched-image recipe + the GB10 memory study. · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session.…
What people said about glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark on X
🚀 New GLM-5.3-Flash default on 4x DGX Spark (TP4) Found 18 GiB of attention + MLP weights still sitting in bf16, read on every single step. Quantized them to NVFP4. ⚡ Per-step time: 67.6ms → 57ms 🧠 500K ctx, fp8 KV, 3.53M token pool 📈 vs our previous GLM lane: Structure +30% | Math +29% | Prose +27% Counting +21%…
— @Tech2Wild, 1 h ago · 16 likes · see the post
Right Now I will swapping the Redhat with the OFFICIAL NVIDIA NVFP4 Checkpoint and see how it holds up. Right now I have not had ANY issues with RedHat and the checkpoint size is smaller so MORE KV Pool.
— @Tech2Wild, 1 h ago · 4 likes · see the post
The community is working on TP4. We moving. Thanks @Tech2Wild
— @TechMDAI, 1 h ago · 4 likes · see the post
Alternatives to glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark
- StepFun Open Platform — Step API · Stable · High-Performance · Easy Integration. Leading models and tools to accelerate the deployment of your…
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- nimble — Code and info
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark in numbers
- Rank on the AI Radar: #67 of 1464
- Shared in 6 posts by 2 accounts: @Tech2Wild, @TechMDAI
- 4.8k views on those posts
- First seen 1 h ago, last shared 1 h ago
- Pricing seen by Jev: open source
- Market: AI infra & evals
FAQ
What is glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark?
GLM-5.3-Flash (320B MoE) at TP4 across four DGX Sparks, same day as the model drop: 36 tok/s, 1.26M-token FP8 KV pool, 262K context, MTP spec decode. First TP4 glm5_next outside B200 hardware. Full... It was first shared on X 1 h ago and is ranked #67 on the AI Radar.
Is glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark free?
It is open source.
Who shared glm-5.3-flash-nvfp4-1m-kv-4x-dgx-spark?
2 accounts on X, including @Tech2Wild, @TechMDAI, in 6 posts totalling 4.8k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.7k posts from 4.9k X accounts over the last 14 days, 1.5k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 17:01 UTC. Full method.