AI Radar
Support
LiveUpdated 2026-09-20 17:01 UTC

HerHealthEval tests clinical LLMs across English, French, Arabic in six registers.

HerHealthEval tests clinical LLMs across English, French, Arabic in six registers. Language-asymmetric supervision:…

This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.

HerHealthEval tests clinical LLMs across English, French, Arabic in six registers. Language-asymmetric supervision: 0.994 under-triage in FR and AR. Source-derived invariant labels: drops to 0.572 and 0.558. Aggregate accuracy hides it. https://arxiv.org/abs/2609.20684

Posted by Alex Zhavoronkov, PhD (aka Aleksandrs Zavoronkovs) (43.8k followers) 1 h ago · 3 likes · 372 views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: arXiv.org.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.7k posts from 4.9k X accounts over the last 14 days, 1.5k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 17:01 UTC. Full method.