qwen3.8-flash-next-mlx-4bit-mtp
qwen3.8-flash-next-mlx-4bit-mtp is We’re on a journey to advance and democratize artificial intelligence through open source and open science.. It is ranked #1013 on the AI Radar, in Open & local models, first seen 9 days ago and shared in 3 posts (42.5k views).
Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP · Hugging Face Hugging Face Log In Sign Up <|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%-…
What people said about qwen3.8-flash-next-mlx-4bit-mtp on X
Qwen 3.8 Next Flash 4bit MTP on my M3U Studio is absolutely flying now. My latest OMP run measured 97.1 tok/s overall decode at 23k context, with ~1,131 tok/s uncached PP. PR opened on oMLX for this. https://huggingface.co/Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP
— @ashxhart, 5 days ago · 130 likes · see the post
MLX and Qwen 3.8 Next are getting some love, and they are using one of my flavours 😎 https://huggingface.co/Vontra/Qwen3.8-Flash-Next-MLX-4bit-MTP I have been tuning this since early this week and am currently sitting at 91.24 tok/s repeatedly, but at ~3.4k context with greedy sampling and a 512-token thinking cap.…
— @ashxhart, 9 days ago · 54 likes · see the post
My Qwen 3.6 35b a3b hits around 85tks and that’s a lot smaller than Qwen 3.8 Flash Next. The pace at which local Ai is advancing is astonishing. Moreover, the AI community on X is hands down the best 🫶🏼
— @ashxhart, 5 days ago · 31 likes · see the post
Alternatives to qwen3.8-flash-next-mlx-4bit-mtp
- Pirate Face — The un-bannable, checksum-verified mirror for open models. Claim your handle before a squatter does.
- atomic.chat — Run AI models locally ->
- deepseek-v4.1-flash-exl3-2x-dgx-sparks — DeepSeek v4.1 Flash EXL3 2.9 bpw for 2x DGX Sparks - MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
- bespoke-nimble-9b — We’re on a journey to advance and democratize artificial intelligence through open source and open science.
- Cactus Compute — It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS,…
- orcabonsai-27b-uncensored — Open source
qwen3.8-flash-next-mlx-4bit-mtp in numbers
- Rank on the AI Radar: #1013 of 1190
- Shared in 3 posts by 1 account: @ashxhart
- 42.5k views on those posts
- First seen 9 days ago, last shared 5 days ago
- Pricing seen by Jev: open source
- Market: Open & local models
FAQ
What is qwen3.8-flash-next-mlx-4bit-mtp?
We’re on a journey to advance and democratize artificial intelligence through open source and open science. It was first shared on X 9 days ago and is ranked #1013 on the AI Radar.
Is qwen3.8-flash-next-mlx-4bit-mtp free?
It is open source.
Who shared qwen3.8-flash-next-mlx-4bit-mtp?
1 account on X, including @ashxhart, in 3 posts totalling 42.5k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.