AI Radar
Support
LiveUpdated 2026-09-21 13:21 UTC

重新部署并测试了 Qwen3.8-Flash-Next 的 OrcaRouter Uncensored NVFP4 版本。

重新部署并测试了 Qwen3.8-Flash-Next 的 OrcaRouter Uncensored NVFP4 版本。 只有一张 RTX PRO 6000 Blackwell 96GB,功耗限制在 350W。后端使用…重新部署并测试了 Qwen3.8-Flash-Next 的 OrcaRouter Uncensored NVFP4 版本。 只有一张 RTX PRO 6000 Blackwell 96GB,功耗限制在 350W。后端使用…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.

重新部署并测试了 Qwen3.8-Flash-Next 的 OrcaRouter Uncensored NVFP4 版本。 只有一张 RTX PRO 6000 Blackwell 96GB,功耗限制在 350W。后端使用 Pennyroyal / SGLang,开启 FP8 KV Cache、PLE 系统内存卸载和 MTP 推测解码,主要给多个 subagent 提供本地推理服务。 最新 sparkDash 实测成绩很猛。 单路 Prose Decode 达到 233.7 tok/s。 2 路并发总吞吐 426.1 tok/s。 4 路达到 657.8 tok/s。 8 路达到 794.9 tok/s。 16 路总吞吐突破 1064.1 tok/s,平均每路仍有 69.4 tok/s。 Prefill 在 8K 到 128K 上下文范围内,基本维持每秒 1.05 万到 1.13 万 tokens。读取 128K 输入只需约 12 秒就能输出首个 token。 当前配置支持 16 个请求同时执行,共享约 100 万 token KV 池,单请求上下文上限为 262,144 tokens。 一张限制在 350W 的 PRO 6000,已经可以同时为十几个本地 Agent 提供接近云端的响应速度。 本地多 Agent 推理,开始真正进入实用阶段了。

Posted by Wei (6.8k followers) 1 h ago · 4 likes · 514 views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 38.5k posts from 5k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 13:21 UTC. Full method.