https://benchmarks.bespokelabs.ai/autoresearchexam/
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
https://benchmarks.bespokelabs.ai/autoresearchexam/ Grok 4.7 added to our live leaderboard for AutoResearchExam. It jumps to position no3 at 30Min research, drops to 4, then again at no3 at 4h, landing at position no4 after 24h of AutoResearch. Pretty good performance and cost tradeoff overall!
Posted by Alex Dimakis (25.1k followers) 1 h ago · 8 likes · 407 views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: AutoResearchExam.
More AI work like this
- Anthropic confirms upcoming models: — @StatsWire
- A necessary condition for a good coding model: it should speak human, not silicon, even… — @sheriyuo
- This website helps you master Claude Opus 5.5 and use it for coding, AI agents,… — @shushant_l
- Tried a new stealth model, Space Bunny. Gave it 1 prompt: make a voxel pagoda garden in… — @superalesha
- Reminds me of this... SpaceXAI literally put this in writing 😂 — @XFreeze
- Grok 4.7 works really well with Blender. I’ve been using them to create these rough… — @cb_doge
- Grok 4.7: Better Self-Testing, but Progress Comes at a Cost — @ZhihuFrontier
- Big news: Claude Opus 5.5 (Max) by @AnthropicAI just topped #1 in Code Arena: WebDev… — @arena
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.9k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 05:42 UTC. Full method.