测试了 4 个在近期发布的模型的鹈鹕骑车:Step 5、Grok 4.7、MiMo V2.6 Flash、MiMo V2.6 Pro
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
测试了 4 个在近期发布的模型的鹈鹕骑车:Step 5、Grok 4.7、MiMo V2.6 Flash、MiMo V2.6 Pro 两个对照组分别是:Gemini 3.8 Flash(左上)、SWE-2(右下) -绝大部分用时3-5min,MiMo V2.6 Pro 用时 10min,唯有 Grok 4.7 用时 37min; -Gemini 3.8 Flash 明显胜出,其他各有优劣 鹈鹕骑车只是个娱乐性的测试,不能全面反应模型的能力,但是 Grok 4.7 做了 40 分钟就交出来这一坨让我实在有点......
Posted by FeiZ (8.2k followers) 1 h ago · 12 likes · 1.7k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- Grok 4.7 is now officially available in Open Design. — @OpenDesignHQ
- GLM-5.3-FlashX is now officially available in Open Design. — @OpenDesignHQ
- Today's boutta be a movie — @vikktorrrre
- GROK 4.7 xHIGH JUST ENTERED THE TOP AI MODEL RACE. — @ZentrixHQ
- We tested GLM 5.3 and Grok 4.7 with the same prompts to build immersive 3D temple… — @OpenDesignHQ
- Qwen 4 is coming soon!! — @ChrisGPT
- This day starts amazing: Qwen 4 announced, coming soon. Even our beloved 27b. — @kimmonismus
- 🚨 GPT-6 Sol is showing up in the served model logs — @Mr_Salio
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 41.4k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 11:47 UTC. Full method.