"HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing"
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
"HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing" Xiaomi dropped their new attention mechanism for their upcoming V3 model. Their new HySparse2 makes long-context agents much cheaper by sharing KV caches across layers instead of repeatedly processing and storing the entire context. With two-level KV sharing and token-level sparse attention, HySparse2 can stop prefill halfway through the model while still retrieving the right tokens from long contexts. At 1M context, it cuts prefill FLOPs by 2.92x vs HySparse and shrinks the KV cache from 6.72GB to 2.69GB, improving long-contex
Posted by alphaXiv (56.3k followers) 2 h ago · 13 likes · 1.3k views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: alphaXiv.
More AI work like this
- How does this only have 21,000 views in 8 days? — @WR4NYGov
- HappyWorld-Bench — @HuggingPapers
- Nice to see AlphaFold and other bio ML models being used in Anthropic's novel enzyme… — @JMateosGarcia
- @Sauers_ — @Sauers_
- Contrastive Language Models — @_akhaliq
- Impressive paper from Salesforce. — @dair_ai
- Yukon: Make Research Open & Multiplayer Again. — @sreeramkannan
- Must-read paper from Google on self-improving agent harnesses. — @omarsar0
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 46k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 11:00 UTC. Full method.