Terminal-Bench 4.0 is now live on Vals.
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Terminal-Bench 4.0 is now live on Vals. It is a benchmark of 66 new tasks that ask an agent to do real terminal work from start to finish, like shipping a working service, proving a theorem, training a GPU kernel, or writing a forensic report. The median task is estimated at 4 hours of expert work, and grading is strict, so a model either clears the full verifier suite or scores nothing. Every score is avg@3: three pass@1 runs, averaged.
Posted by Vals AI (21.1k followers) 2 days ago · 151 likes · 12.4k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- gpt 5.6 terra = gpt 6 astra — @notjazii
- wow — @HarshithLucky3
- 🚨 Gemini 4 Pro is absolutely cooking — @Mr_Salio
- anthropic is about to mog everyone with new models — @notjazii
- woke people still don’t want to accept AIs good potential, while AI be like to them🖕🏻 — @SciTechera
- @AdamHoltererer — @AdamHoltererer
- GPT 6 Astra Pro builds OCD simulator: — @AdamHoltererer
- july 20, 2027. claude 6 mendel waking up in anthropic wet lab — @dejavucoder
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:45 UTC. Full method.