Jev is a week old and has 31 open reproductions. Most comparisons so far run a few…
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Jev is a week old and has 31 open reproductions. Most comparisons so far run a few hundred decisions. interesting new benchmark: Decision Index 0.1 - it runs 132,422: every repro + Jev on 37 benchmarks. No truncation, no prompt tuning, unanswered equals wrong. Jev still tops it at 59.5 👀 https://huggingface.co/spaces/multimodalart/jev-decision-index
Posted by Victor M (30.8k followers) 1 h ago · 47 likes · 1.5k views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: jev-decision-index.
More AI work like this
- openai is preparing to release new bot "aeon" — @notjazii
- Grok 4.7 High vs Xiaomi MiMo-V2.6 ! — @YouWareAI
- Mimo-v2.6-flash on my one-shot side-scroller. — @NeoAIForecast
- BREAKING: xAI just shipped Grok 4.7 and it outperforms GPT-5.6 Sol Max at less than half… — @guilleflorvs
- Confirmed by Alibaba: Qwen 4 27B is coming — @AestherML
- 🚨 OPUS 5.5 Leaks — @Priyannkaaaa
- 阿里的Qwen4 已经在训练中,即将正式发布: — @mylifcc
- Grok4.7第一波实测来了,Grok 4.7 (high in Grok build)和GPT-5.6 Sol(high in codex) — @xueyu1125
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 41.2k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 09:09 UTC. Full method.