AI Radar
Support
LiveUpdated 2026-09-19 18:08 UTC

Onur Solmaz blog

Onur Solmaz blog — The small KV cache will also do something good for local inference, but I need more time…

Onur Solmaz blog is The small KV cache will also do something good for local inference, but I need more time to calculate how much. Take these with a grain of salt. Give your agent. It is ranked #167 on the AI Radar, in AI infra & evals, first seen 9 days ago and shared in 4 posts (268.7k views).

Visit solmaz.io

What Onur Solmaz blog says about itself

Estimate LLM throughput from hardware and model limits, with formulas and interactive examples for decode and prefill.

Theoretical Upper Bounds for LLM Throughput | Onur Solmaz blog Onur Solmaz blog Post Theoretical Upper Bounds for LLM Throughput / 2026 / 07 / 14 • /2026/08/28 • 54 min read Copy markdown Copy link Paper view On this page Four questions Memory power Two numbers and their product Feasibility theorem A caveat on the metric A first decode bound A toy decoder Capacity limit Bandwidth limit Memory-power decode bound Scope of the bound Missing traffic Universal resource bound Bytes per token Memory-power bound as a corollary KV-aware bound Two context lengths Model quantities Memory-fit batch Aggregate and per-session ceilings Usable-batch…

What people said about Onur Solmaz blog on X

DeepSeek V4.1 Flash is NOT an LLM It is a hyper-efficient, invasive species designed to displace every other LLM, similar to how Döner displaced all other fast food in Germany As if V4 wasn't cheap and efficient enough, they made it a LOT cheaper 2.3x cheaper cached input tokens???? 1.5x cheaper uncached input and…

@onusoz, 9 days ago · 1.3k likes · see the post

DeepSeek V4.1 Flash weights are out! ⚡️⚡️⚡️ I have good news and bad news Don't be fooled by "V4".1, this is a different architecture V4 was 284B total, 13B active V4.1 has a 552B MoE backbone + 196B engram conditional-memory parameters with 16B active parameters Bad news first. A single DGX Spark will likely not be…

@onusoz, 9 days ago · 139 likes · see the post

RIP CLAUDE.md 24.02.2025 - 18.09.2026 (572 days) Anthropic has finally yielded, now you can use just AGENTS.md for agent instructions This was my most visited webpage at some point. So much time and money wasted: https://solmaz.io/log/2025/09/08/claude-md-agents-md-migration-guide

@onusoz, 8 h ago · 35 likes · see the post

Alternatives to Onur Solmaz blog

Onur Solmaz blog in numbers

FAQ

What is Onur Solmaz blog?

The small KV cache will also do something good for local inference, but I need more time to calculate how much. Take these with a grain of salt. Give your agent It was first shared on X 9 days ago and is ranked #167 on the AI Radar.

Is Onur Solmaz blog free?

Pricing is not stated on the page we read.

Who shared Onur Solmaz blog?

2 accounts on X, including @onusoz, @TechMDAI, in 4 posts totalling 268.7k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.