cuda.fast
cuda.fast is Make Qwen 3.8 125B-A6B decode faster with the MTP head on CUDA and MLX, and follow every improvement on the live cuda.fast leaderboard.. It is ranked #1052 on the AI Radar, in AI infra & evals, first seen 9 days ago and shared in 1 post (9k views).
Visit yukon-challenges-ui-git-qwen38flashtest-eigen-labs.vercel.app
cuda.fast — Qwen 3.8 125B-A6B, faster with speculative decode Yukon cuda .fast All challenges Leaderboard Qwen 3.8 125B-A6B runs 18.8% 0.0% 18.8% faster on cuda.fast . How the score works The official score is the single-stream composite: prefill gain to the 0.25 power times decode gain to the 0.75 power, so decode carries three quarters of the weight. A serial run is 1.0, and everything above it is real speculative-decode gain. benchd computes the score and seals it. cuda.fast mlx.fast How it works Participate Record By model Lin Log cuda.fast 18.8% Improvement History Leaderboard current record 18.8% composite speed increase over the…
What people said about cuda.fast on X
ATTENTION ALL DGX SPARK OWNERS....NOW IS OUR TIME TO SHINE! The CUDA vs MLX Qwen 3.8 125B A6B challenge has begun! Who has the best and most active community that can push this model to the fastest speed? https://yukon-challenges-ui-git-qwen38flashtest-eigen-labs.vercel.app/qwen38flashtest
— @GumbiiDigital, 9 days ago · 57 likes · see the post
Alternatives to cuda.fast
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
cuda.fast in numbers
- Rank on the AI Radar: #1052 of 1190
- Shared in 1 post by 1 account: @GumbiiDigital
- 9k views on those posts
- First seen 9 days ago, last shared 9 days ago
- Pricing seen by Jev: not stated
- Market: AI infra & evals
FAQ
What is cuda.fast?
Make Qwen 3.8 125B-A6B decode faster with the MTP head on CUDA and MLX, and follow every improvement on the live cuda.fast leaderboard. It was first shared on X 9 days ago and is ranked #1052 on the AI Radar.
Is cuda.fast free?
Pricing is not stated on the page we read.
Who shared cuda.fast?
1 account on X, including @GumbiiDigital, in 1 post totalling 9k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.