🚨OpenAI just published a mental-health benchmark where several frontier AI models…
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
🚨OpenAI just published a mental-health benchmark where several frontier AI models scored higher than manually written responses from mental-health experts. GPT-6 Astra: 57.3 Claude Opus 5.5: 52.4 Expert responses: 38.5 The interesting part: clinicians actually incurred fewer penalties. OpenAI says they scored lower largely because they answered like they would in person, extremely briefly, often with a single question or statement.
Posted by Choblin (3k followers) 1 h ago · 10 likes · 1.5k views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: openai.com.
More AI work like this
- Opus 5.5 built this 3D arcade shooter with Higgsfield. — @higgsfield_ai
- Opus 5.5 made little Claudies — @blueemi99
- GPT-6 Astra landed on every planet and moon in KSP with a surface (14 worlds) and flew a… — @ValsAI
- Check out side-by-side outputs generated on Arena from Claude Opus 5.5 by @AnthropicAI… — @arena
- Ok one more fun Opus 5.5 demo: — @petergyang
- Across 360 robotics trials, Opus 5.5 averages under $1 per trial, scoring 1.8x as high… — @chooi_jeq
- Millions of people turn to AI to navigate difficult relationships, work through stress,… — @thekaransinghal
- Places to use Jev for faster and cheaper work. — @nextbigfuture
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.1k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 20:32 UTC. Full method.