New benchmark from OpenAI: MentalHealth Bench
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
New benchmark from OpenAI: MentalHealth Bench Interestingly, the human expert responses scored below most models (38.5% for clinician-written replies versus 57.3% for Astra) They explain in a way that clinicians often wrote a short / natural next turn while the rubric rewards covering several useful behaviors in one answer. Quite a striking example of a benchmark score diverging from how an expert might actually converse.
Posted by himanshu (29.1k followers) 1 h ago · 4 likes · 459 views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- Claude Opus 5.5 is the greatest AI model ever released — @AlexFinn
- GLM-5.3 beat Space Bunny Alpha (the new stealth model launched today) on Newton’s cradle… — @rohanpaul_ai
- The token usage chart here is really striking — @charliermarsh
- Anthropic’s guide to using Opus 5.5 has one central rule: stop asking it to think. Eddie… — @davidarngar
- Claude Opus 5.5 recreated an athlete’s long jump in 4D. — @higgsfield_ai
- Claude Opus 5.5 painted the past and future of our world in JavaScript + Higgsfield. — @higgsfield_ai
- Space Bunny performs TERRIBLE in physics 💀 — @aimlapi
- Okay, my first complete benchmark of Jev with my new Decision v1 eval suite on… — @morganlinton
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.6k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 00:57 UTC. Full method.