AI Radar
Support
LiveUpdated 2026-09-19 19:27 UTC

刚在 oMLX 的社区跑分页面看到一组 M5 Ultra 数据,不知道有多真,但数字看起来相当合理。

刚在 oMLX 的社区跑分页面看到一组 M5 Ultra 数据,不知道有多真,但数字看起来相当合理。 Qwen3.8 27B Q4 在 8K 上下文下,生成约 50 tok/s,Prefill 约 1800 tok/s,而且还没有开启…

This is a AI post classified by Jev as Open & local models (a model release), kept by the AI Radar because it carries real work, not commentary.

刚在 oMLX 的社区跑分页面看到一组 M5 Ultra 数据,不知道有多真,但数字看起来相当合理。 Qwen3.8 27B Q4 在 8K 上下文下,生成约 50 tok/s,Prefill 约 1800 tok/s,而且还没有开启 MTP。 如果数据属实,这已经是非常舒服的本地 Agent 速度。50 tok/s 足够实时写代码和调用工具,1800 tok/s 的 Prefill 也能快速读取大型项目和长文档。 后续加上 MTP、ANE Prefill 和更成熟的 oMLX 优化,生成速度还有继续提高的空间。

Posted by Wei (6.8k followers) 2 days ago · 17 likes · 5k views · view the original post on X. Kept by the AI Radar as Open & local models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.3k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 19:27 UTC. Full method.