Full Kimi K3 (2.8T) frontier model running locally on 16× NVIDIA GB10 Spark Desk…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
Full Kimi K3 (2.8T) frontier model running locally on 16× NVIDIA GB10 Spark Desk cluster, Not a data center. - 30 t/s single stream - 87 t/s aggregate at C8 - 136 t/s peak claimed - Prose llama-bench holding 20–23 t/s generate even at long context. You don’t need a hyperscaler rack to run a 2.8T model anymore.
Posted by Md Ismail Šojal 🕷️ (55.6k followers) 1 h ago · 4 likes · 837 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- I built ds-router for Hermes Agent. — @btsouth
- Happy weekend! Let's forget all the troubles in daily life; just go outside to touch… — @alibaba_cloud
- jev as a Judge proves to be a cheaper and more precise alternative to LLM as a judge for… — @sydneyrunkle
- Forked Omarchy’s built-in AI usage plugin and added a full dashboard with token history,… — @btsouth
- BIG MILESTONE — @thepirateface
- I WOKE UP TO MONEY — @michael_chomsky
- Local LLM Tips: — @Hikari_07_jp
- we launched the most comprehensive ai performance engineering repo in the world — @wafer_ai
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 31.6k posts from 4.8k X accounts over the last 14 days, 1.3k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 02:50 UTC. Full method.