Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode,…
This is a AI post classified by Jev as AI infra & evals (a launch), kept by the AI Radar because it carries real work, not commentary.
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
Posted by Inferact (6.4k followers) 1 h ago · 118 likes · 8.5k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- This just in! Take our live, in-person PyTorch Associate Training the day before PyTorch… — @PyTorch
- Want to run GLM 5.3 Flash EXL3 on 2x DGX Spark like I do? My full config is now in the… — @plotarmordev
- One week left on 90% off GLM 5.3 Flash. — @merge_api
- ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700… — @SemiAnalysis_
- very impressive work by @damian_b @a_kirillo and others, building the sandboxing part of… — @eliebakouch
- Qwen3.7-Max and Qwen3.8-Flash from @Alibaba_Qwen are now 40% off on serverless through… — @togethercompute
- Introducing Prime Sandboxes: — @PrimeIntellect
- Model Vault is now available in Canada 🇨🇦 — @cohere
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 20:02 UTC. Full method.