AI Radar
Support
LiveUpdated 2026-09-19 18:08 UTC

LLM Context Benchmark Splash inference engine by @inco_ai running on M5 Max 128GB macOS…

LLM Context Benchmark Splash inference engine by @inco_ai running on M5 Max 128GB macOS 27 in high power mode…LLM Context Benchmark Splash inference engine by @inco_ai running on M5 Max 128GB macOS 27 in high power mode…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.

LLM Context Benchmark Splash inference engine by @inco_ai running on M5 Max 128GB macOS 27 in high power mode (throttling here and there). incoai-internal/Qwen3.8-27B-Splash Hardware: Apple M5 Max, 128.0GB RAM, 18 CPU cores, 40 GPU cores 0.5k pp 894 tg 106 t/s 1k pp 981 tg 84 t/s 2k pp 989 tg 81 t/s 4k pp 917 tg 70 t/s 8k pp 896 tg 80 t/s 16k pp 878 tg 101 t/s 32k pp 805 tg 72 t/s 64k pp 691 tg 93 t/s 128k pp 531 tg 53 t/s 256k pp 388 tg 57 t/s Total generated tokens: 1280 Batch TPS: b1 58 b2 96 b4 144 b8 136 More info in their blog post: https://inco.ai/blog/splash/

Posted by Ivan Fioravanti ᯅ (42.7k followers) 17 h ago · 50 likes · 4.8k views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.