Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core OpenAI models available today, across all effort levels. Through this process as I shared a bit last week, I also decided to create a third eval suite, focused on routine engineering tasks. This current suite, which I'm now calling VulcanBench Frontier, is really, most likely, harder tasks that regular engineers on engineering teams are giving models on a normal day. What I'm testing with this eval suite is how these models do with hard stuff, things you might give a model, bu
Posted by Morgan (46.1k followers) 1 days ago · 432 likes · 81.1k views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: VulcanBench.
More AI work like this
- Confirmed: GPT-6 Sol releasing on Tuesday — @StatsWire
- Google may have cooked hard this time. A new model on Arena, widely believed to be… — @ZentrixHQ
- I can't confirm any of this, but I'm at least hearing the rumors. — @kimmonismus
- Anthropic is skipping Opus 5.2 entirely for a "cross-level" surprise attack on OpenAI's… — @StatsWire
- Jev pod tomorrow — @swyx
- Gemini 4 Pro just made this. — @0x0SojalSec
- Just discovered CancerBench was also popular on Threads lol — @iScienceLuvr
- Opus 5.5 is currently being ghost tested through Sonnet 5 — @MaaSonder
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 38k posts from 4.9k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 07:04 UTC. Full method.