Okay, my first complete benchmark of Jev with my new Decision v1 eval suite on…
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Okay, my first complete benchmark of Jev with my new Decision v1 eval suite on @VulcanBench is done. I'll be honest, I'm still not totally happy with it, definitely needs some improvement. But there's something here, and it's starting to get more interesting to me, so I thought I would share it. As an independent benchmarker, I care less about perfection, and more about the process of learning and iterating, and generating as much signal and insights as I can along the way. I think the key insight I was able to get from this first run on Verdict v1 is that Jev is good at judging how well co
Posted by Morgan (46.1k followers) 1 h ago · 11 likes · 623 views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: VulcanBench.
More AI work like this
- Claude Opus 5.5 is the greatest AI model ever released — @AlexFinn
- GLM-5.3 beat Space Bunny Alpha (the new stealth model launched today) on Newton’s cradle… — @rohanpaul_ai
- The token usage chart here is really striking — @charliermarsh
- Anthropic’s guide to using Opus 5.5 has one central rule: stop asking it to think. Eddie… — @davidarngar
- Claude Opus 5.5 recreated an athlete’s long jump in 4D. — @higgsfield_ai
- Claude Opus 5.5 painted the past and future of our world in JavaScript + Higgsfield. — @higgsfield_ai
- Space Bunny performs TERRIBLE in physics 💀 — @aimlapi
- New benchmark from OpenAI: MentalHealth Bench — @himanshustwts
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.6k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 00:57 UTC. Full method.