BeeLlama v0.4.7 Preview just made DFlash 2 by @inco_ai much more efficient in terms of…
This is a AI post classified by Jev as AI infra & evals (a tool drop), kept by the AI Radar because it carries real work, not commentary.
BeeLlama v0.4.7 Preview just made DFlash 2 by @inco_ai much more efficient in terms of memory consumption. The issue: in llama.cpp, a larger ubatch size greatly improves prefill performance, but also requires more VRAM. What's worse, DFlash pays the same price twice, because the drafter is a separate model. The solution is simple: allow setting a different ubatch size for the drafter. In the example below, the target model uses ub 1024. Keeping it as is while overriding the drafter to ub 128 saves 1.8 GB of VRAM without any negative effects (the drafter is small and doesn't need a large ubat
Posted by Ivan Neustroev (272 followers) 1 h ago · 1 likes · 20 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- I am finally learning how to fine-tune models using @UnslothAI doing a "hello world" test. — @jvr0x
- Nearly three years ago, Rendair started with just one serverless endpoint on Runpod. — @runpod
- Stay on top of model changes, low credits, workspace budgets, and API key limits with… — @OpenRouter
- Me looking at "Apple's M5 Ultra Mac Studio Shines in Local AI Tests" trending on X with… — @ivanfioravanti
- Your users have already generated a better eval dataset than you'd write yourself. — @sentry
- Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high… — @vllm_project
- 🚀 Introducing Luminary, our AI wind tunnel for enterprise model selection. — @MultiverseCompu
- cool how Hugging Face made tokenizers 15.9× faster without changing a single token ID: — @victormustar
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 39.1k posts from 5k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 16:32 UTC. Full method.