AI Radar
Support
LiveUpdated 2026-09-19 18:08 UTC

Splash vs mlx-vlm+MTP, Qwen3.8-27B 4-bit

Splash vs mlx-vlm+MTP, Qwen3.8-27B 4-bit Mac Studio M4 Max 128 GB. Stock splash serve, nothing else on the GPU, 5…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.

Splash vs mlx-vlm+MTP, Qwen3.8-27B 4-bit Mac Studio M4 Max 128 GB. Stock splash serve, nothing else on the GPU, 5 coding prompts, 1,024-token cap, medians: Code, no reasoning: Splash 68.0 · MLX 48.9 Code, reasoning on: 53.7 · 34.7 32K+ thinking run: 33.9 · 28.4 Short prose: 36.3 · 38.3 Splash is 20–40% faster on code on this box. Ties on prose. Their 74 tok/s is an M5 Pro number; M4 Max lands 54–68 depending on reasoning. Two things to know: default 32K output cap and 30-min request timeout will cut off a long generation. @incoai @zhijianliu_

Posted by Wësche (2.9k followers) 15 h ago · 25 likes · 2.5k views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.