Vision tower on DeepSeek-V4-Flash-Vision-Exp is 0.41B params. The whole model is 304.65B.
This is a AI post classified by Jev as Open & local models (a model release), kept by the AI Radar because it carries real work, not commentary.
Vision tower on DeepSeek-V4-Flash-Vision-Exp is 0.41B params. The whole model is 304.65B. 0.135 percent. The eyes are under a GB even at bf16. 97.3 percent of the checkpoint is routed experts, shipped as native fp4. Dense parts fp8. So a 304B model lands at 167.8GB on disk, 4.41 bits per weight before anyone re-quants it. That is the local wall. 167.8GB does not fit a 128GB Strix Halo UMA, and that is before a single token of KV cache. Unsloth has a GGUF out, have not run it yet, so no tok/s claim from me. Config math only. Config also lists one KV head, alternating compress ratios of 4 a
Posted by 코지베어 🐻 CozyBear (1.2k followers) 1 h ago · 0 likes · 5 views · view the original post on X. Kept by the AI Radar as Open & local models.
More AI work like this
- Local Ai is about to enter the stratosphere 🤯🚀 — @ashxhart
- The Easiest way to start with Local AI Today? ODS, which is completely free (open-source… — @TheAhmadOsman
- Meet Sharp-Spark-X2.5-4B-GGUF: a 4B parameter text generation model built for speed and… — @HuggingModels
- We rent our intelligence. — @MiaAI_lab
- Today at Snapdragon Summit, we’re announcing 1-bit Bonsai vision-language model running… — @PrismML
- Text to image just got a whole lot faster. Qwen Image 2.1 in GGUF format is here, ready… — @HuggingModels
- dear rtx 3060, 3070, 3080 and every 8gb to 12gb rtx owner, your gpu can run 27b dense… — @sudoingX
- Alright, the weird loop it got into fetching the information seems to be related to omp. — @jvr0x
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.3k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 22:05 UTC. Full method.