ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700…
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user, 56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3. As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.
Posted by SemiAnalysis (166.6k followers) 2 h ago · 532 likes · 61.8k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Your AI agents are generating a ton of operational data. — @AlphaSignalAI
- Most people who use AI and invest in AI Infrastructure operate at a high abstraction… — @siliconcodesign
- This just in! Take our live, in-person PyTorch Associate Training the day before PyTorch… — @PyTorch
- Want to run GLM 5.3 Flash EXL3 on 2x DGX Spark like I do? My full config is now in the… — @plotarmordev
- One week left on 90% off GLM 5.3 Flash. — @merge_api
- very impressive work by @damian_b @a_kirillo and others, building the sandboxing part of… — @eliebakouch
- Qwen3.7-Max and Qwen3.8-Flash from @Alibaba_Qwen are now 40% off on serverless through… — @togethercompute
- Introducing Prime Sandboxes: — @PrimeIntellect
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.1k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 20:32 UTC. Full method.