AI Radar
Support
LiveUpdated 2026-09-23 20:32 UTC

OpenAI's new MentalHealthBench puts GPT-6 Astra at 57.3, versus GPT-4o's 32.1, on…

OpenAI's new MentalHealthBench puts GPT-6 Astra at 57.3, versus GPT-4o's 32.1, on realistic mental health…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

OpenAI's new MentalHealthBench puts GPT-6 Astra at 57.3, versus GPT-4o's 32.1, on realistic mental health conversations. Most mental health AI evaluations have centered on emergencies and broad safety criteria, leaving everyday and ambiguous conversations much less measured. So OpenAI co-created this benchmark with more than 80 licensed psychologists and psychiatrists from 22 countries, spanning 19 languages and nearly 20 subspecialties. Each synthetic conversation gets a custom expert rubric covering behaviors such as seeking context, preserving user agency, safety, and appropriate guidanc

Posted by Rohan Paul (157.9k followers) 1 h ago · 11 likes · 2.3k views · view the original post on X. Kept by the AI Radar as Frontier models. Tools mentioned: openai.com.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.1k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 20:32 UTC. Full method.