Step 5 Preview is putting some strong results on coding and agentic benchmarks.
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Step 5 Preview is putting some strong results on coding and agentic benchmarks. on DeepSWE v1.1 it scores 67.7, ahead of Kimi-K3 and glm-5.3. Terminal-Bench v4 comes in at 33.3, more than 2.5x than Kimi K3's score, and on GDPval-AA it reaches 1571, more than Kimi-K3. some other benchmarks: > ProgramBench: 80.5 > StepCodeBench: 49 > Agent's Last Exam: 29.5 in one 24-hour test, Step 5 Preview kept iterating on a GPU kernel for ~22 hours and reached 508 TFLOPS, ahead of Opus 5 at 493. it also matched Opus 5 on a 24-hour automated post-training task while using fewer annotator tokens. for a
Posted by Ananth (3.3k followers) 1 h ago · 12 likes · 704 views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- 🚨 GROK 4.7 Drops Today — @Mr_Salio
- Anthropic is gearing up to compete with OpenAI at the forefront of model releases, and… — @testingcatalog
- aaaand it's gone, hope it's coming back later today, this will be a packed release week… — @cheatyyyy
- 🚨 Grok 4.7 is in the OpenCode Zen model catalogue — @LuminaBench
- Gemini is now #14 in Artificial Analysis Intelligence Benchmark — @MaaSonder
- 🚨 Grok 4.7 is in the OpenCode Go model catalogue — @LuminaBench
- What if I told you Grok 4.7 will release today, early afternoon? — @ns123abc
- Grok 4.7 is live on OpenCode Zen! — @cheatyyyy
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 38.2k posts from 4.9k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 10:13 UTC. Full method.