Step 5 Preview: Coding Strength, Uneven Capabilities


This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Step 5 Preview: Coding Strength, Uneven Capabilities @StepFun_ai's Step 5 Preview looks more promising as a coding model than an all-rounder. Zhihu contributor Kitt在进化 points to strong terminal and scientific coding results, alongside weaker knowledge and document performance. Look beyond the overall ranking The author cautions against treating Artificial Analysis's headline score as a precise measure of capability. The individual tests are more revealing. In the results cited, Step 5 Preview scores 33% on Terminal-Bench 4.0, versus 27% for DeepSeek V4.1 Flash. It also reaches 59% on SciCode,
Posted by Zhihu Frontier (12.2k followers) 1 h ago · 4 likes · 434 views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- Gemini 4 Pro benchmarks just dropped — @MaaSonder
- @Luna11054 — @Luna11054
- Fable 5.2 is pure JS thats it one shot — @chetaslua
- 🚨 Fable 5.2 Leak: Beats GPT-6 Astra — @Priyannkaaaa
- Best way to test Step-5-preview is with comparison. — @ItsmeAjayKV
- Step 5 Preview — @HarshithLucky3
- 🚨 Step 5 Preview is actually pretty cracked — @LuminaBench
- 🚨 StepFun just dropped Step 5 Preview — @LuminaBench
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.2k posts from 4.8k X accounts over the last 14 days, 1.4k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 10:41 UTC. Full method.