mlx-serve
mlx-serve is Qwen3.8 Flash-Next already has MTP and a large mlx-serve-specific kernel stack, but sampled creative requests could only use exact speculative acceptance. A matched llmprobe creative/thinking-off s.... It is ranked #363 on the AI Radar, in AI infra & evals, first seen 4 days ago and shared in 2 posts (4.2k views).
Port sampled MTP verifiers and paired Qwen3.8 routed gate/up by davidtai · Pull Request #427 · ddalcu/mlx-serve · GitHub Skip to content Navigation Menu Sign in Appearance settings Search / Sign in Sign up Appearance settings You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} ddalcu / mlx-serve Public Uh oh! There was an error while loading. Please reload this page . Notifications You must be signed in to change notification settings Fork…
What people said about mlx-serve on X
So we can stay over 100tps at 128k context and get ~70tps on Qwen 3.8 Flash Next at 1 million context. I wanted to see how well the relaxed verifiers worked on super large contexts so I ported them to MLXServe. MLXServe recently added extended context support to 1M using a combination of tight memory management and…
— @TheDavidTai, 4 days ago · 20 likes · see the post
My relaxed verifier work is also here along with some bug fixes I found that the team has fixed.
— @TheDavidTai, 3 days ago · 14 likes · see the post
Alternatives to mlx-serve
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
mlx-serve in numbers
- Rank on the AI Radar: #363 of 1192
- Shared in 2 posts by 1 account: @TheDavidTai
- 4.2k views on those posts
- First seen 4 days ago, last shared 3 days ago
- Pricing seen by Jev: open source
- Market: AI infra & evals
FAQ
What is mlx-serve?
Qwen3.8 Flash-Next already has MTP and a large mlx-serve-specific kernel stack, but sampled creative requests could only use exact speculative acceptance. A matched llmprobe creative/thinking-off s... It was first shared on X 4 days ago and is ranked #363 on the AI Radar.
Is mlx-serve free?
It is open source.
Who shared mlx-serve?
1 account on X, including @TheDavidTai, in 2 posts totalling 4.2k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:40 UTC. Full method.