重新部署并测试了 Qwen3.8-Flash-Next 的 OrcaRouter Uncensored NVFP4 版本。

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
重新部署并测试了 Qwen3.8-Flash-Next 的 OrcaRouter Uncensored NVFP4 版本。 只有一张 RTX PRO 6000 Blackwell 96GB,功耗限制在 350W。后端使用 Pennyroyal / SGLang,开启 FP8 KV Cache、PLE 系统内存卸载和 MTP 推测解码,主要给多个 subagent 提供本地推理服务。 最新 sparkDash 实测成绩很猛。 单路 Prose Decode 达到 233.7 tok/s。 2 路并发总吞吐 426.1 tok/s。 4 路达到 657.8 tok/s。 8 路达到 794.9 tok/s。 16 路总吞吐突破 1064.1 tok/s,平均每路仍有 69.4 tok/s。 Prefill 在 8K 到 128K 上下文范围内,基本维持每秒 1.05 万到 1.13 万 tokens。读取 128K 输入只需约 12 秒就能输出首个 token。 当前配置支持 16 个请求同时执行,共享约 100 万 token KV 池,单请求上下文上限为 262,144 tokens。 一张限制在 350W 的 PRO 6000,已经可以同时为十几个本地 Agent 提供接近云端的响应速度。 本地多 Agent 推理,开始真正进入实用阶段了。
Posted by Wei (6.8k followers) 1 h ago · 4 likes · 514 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Access all top AI models in one place with Alibaba Cloud Model Studio. Build with… — @alibaba_cloud
- 🚨 A WEBSITE JUST DROPPED 194 THINGS YOU CAN BUILD WITH JEV. NOT A DEMO. A LIVE SHOWCASE. — @Priyannkaaaa
- How fast is Qwen-image-2.1 on a single NVIDIA DGX Spark. — @MichaelGannotti
- Anthropic has rebranded Workbench as Playground! It has been in development for quite a… — @testingcatalog
- At HUAWEI CONNECT 2026, Huawei unveiled its UnifiedBus-powered computing architecture,… — @Huawei
- what did they feed this model... — @zainhas
- Looking at the ASUS ProArt GR1X, powered by the RTX Spark, as expected it's lacking the… — @MiaAI_lab
- $ORBIO — http://ORBIO.SO — @RobinhoodAlphas
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 38.5k posts from 5k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 13:21 UTC. Full method.