Onur Solmaz blog
Onur Solmaz blog is The small KV cache will also do something good for local inference, but I need more time to calculate how much. Take these with a grain of salt. Give your agent. It is ranked #167 on the AI Radar, in AI infra & evals, first seen 9 days ago and shared in 4 posts (268.7k views).
What Onur Solmaz blog says about itself
Estimate LLM throughput from hardware and model limits, with formulas and interactive examples for decode and prefill.
Theoretical Upper Bounds for LLM Throughput | Onur Solmaz blog Onur Solmaz blog Post Theoretical Upper Bounds for LLM Throughput / 2026 / 07 / 14 • /2026/08/28 • 54 min read Copy markdown Copy link Paper view On this page Four questions Memory power Two numbers and their product Feasibility theorem A caveat on the metric A first decode bound A toy decoder Capacity limit Bandwidth limit Memory-power decode bound Scope of the bound Missing traffic Universal resource bound Bytes per token Memory-power bound as a corollary KV-aware bound Two context lengths Model quantities Memory-fit batch Aggregate and per-session ceilings Usable-batch…
What people said about Onur Solmaz blog on X
DeepSeek V4.1 Flash is NOT an LLM It is a hyper-efficient, invasive species designed to displace every other LLM, similar to how Döner displaced all other fast food in Germany As if V4 wasn't cheap and efficient enough, they made it a LOT cheaper 2.3x cheaper cached input tokens???? 1.5x cheaper uncached input and…
— @onusoz, 9 days ago · 1.3k likes · see the post
DeepSeek V4.1 Flash weights are out! ⚡️⚡️⚡️ I have good news and bad news Don't be fooled by "V4".1, this is a different architecture V4 was 284B total, 13B active V4.1 has a 552B MoE backbone + 196B engram conditional-memory parameters with 16B active parameters Bad news first. A single DGX Spark will likely not be…
— @onusoz, 9 days ago · 139 likes · see the post
RIP CLAUDE.md 24.02.2025 - 18.09.2026 (572 days) Anthropic has finally yielded, now you can use just AGENTS.md for agent instructions This was my most visited webpage at some point. So much time and money wasted: https://solmaz.io/log/2025/09/08/claude-md-agents-md-migration-guide
— @onusoz, 8 h ago · 35 likes · see the post
Alternatives to Onur Solmaz blog
- Ling-3.0-flash-Fin — 🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
- classifier.dev — now outperforms jev and is free
- Jev API, Pricing & Playground — Jev is now available on the @vercel AI Gateway
- Darkbloom — Private AI inference through hardware-attested Apple Silicon providers. Your prompts stay encrypted, your data stays…
- glm-5.3-flash-exl3-2x-dgx-sparks — GLM-5.3 Flash EXL3 for 2-4x DGX Sparks. Contribute to MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks development by…
- Intelligence Benchmarking — Detailed intelligence benchmarking methodology for LLM quality evaluations.
Onur Solmaz blog in numbers
- Rank on the AI Radar: #167 of 1190
- Shared in 4 posts by 2 accounts: @onusoz, @TechMDAI
- 268.7k views on those posts
- First seen 9 days ago, last shared 1 h ago
- Pricing seen by Jev: not stated
- Market: AI infra & evals
FAQ
What is Onur Solmaz blog?
The small KV cache will also do something good for local inference, but I need more time to calculate how much. Take these with a grain of salt. Give your agent It was first shared on X 9 days ago and is ranked #167 on the AI Radar.
Is Onur Solmaz blog free?
Pricing is not stated on the page we read.
Who shared Onur Solmaz blog?
2 accounts on X, including @onusoz, @TechMDAI, in 4 posts totalling 268.7k views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.