AI Radar
Support
LiveUpdated 2026-09-19 18:45 UTC

Terminal-Bench 4.0 is now live on Vals.

Terminal-Bench 4.0 is now live on Vals. It is a benchmark of 66 new tasks that ask an agent to do real terminal work…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Terminal-Bench 4.0 is now live on Vals. It is a benchmark of 66 new tasks that ask an agent to do real terminal work from start to finish, like shipping a working service, proving a theorem, training a GPU kernel, or writing a forensic report. The median task is estimated at 4 hours of expert work, and grading is strict, so a model either clears the full verifier suite or scores nothing. Every score is avg@3: three pass@1 runs, averaged.

Posted by Vals AI (21.1k followers) 2 days ago · 151 likes · 12.4k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:45 UTC. Full method.