AI Radar
Support
LiveUpdated 2026-09-21 21:52 UTC

Quick update on the Terminal-Bench mismatch, Artificial Analysis isn’t just copying…

Quick update on the Terminal-Bench mismatch, Artificial Analysis isn’t just copying xAI’s 38% they independently rerun…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Quick update on the Terminal-Bench mismatch, Artificial Analysis isn’t just copying xAI’s 38% they independently rerun the full 66-task Terminal-Bench 4.0, using mini-swe-agent with pass@1 averaged over 3 runs. They also separately test Grok 4.7 with its native Grok Build harness, where they get -33%. I was wondering the same thing as Dan!

Posted by Chris (66.5k followers) 1 h ago · 67 likes · 7.4k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.3k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 21:52 UTC. Full method.