AI Radar
Support
LiveUpdated 2026-09-21 21:52 UTC

Legal work saw one of Grok 4.7’s biggest jumps.

Legal work saw one of Grok 4.7’s biggest jumps. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while…Legal work saw one of Grok 4.7’s biggest jumps. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while…Legal work saw one of Grok 4.7’s biggest jumps. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while…Legal work saw one of Grok 4.7’s biggest jumps. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Legal work saw one of Grok 4.7’s biggest jumps. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while GPT-5.6 Sol Max scored 2.5% and Fable 5.1 Max scored 6.7% Also, Long-running terminal work nearly doubled. Grok 4.7 reached 38.0% on Terminal-Bench 4.0, versus 20.3% for Grok 4.6, suggesting the larger improvement is in agents that must keep working through extended computer tasks. And SpaceXAI made those gains without raising the headline token price. Grok 4.7 remains at $2/$6 per 1M input/output tokens.

Posted by Rohan Paul (157.8k followers) 3 h ago · 16 likes · 2.4k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.3k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 21:52 UTC. Full method.