HerHealthEval tests clinical LLMs across English, French, Arabic in six registers.
This is a AI post classified by Jev as AI research (research), kept by the AI Radar because it carries real work, not commentary.
HerHealthEval tests clinical LLMs across English, French, Arabic in six registers. Language-asymmetric supervision: 0.994 under-triage in FR and AR. Source-derived invariant labels: drops to 0.572 and 0.558. Aggregate accuracy hides it. https://arxiv.org/abs/2609.20684
Posted by Alex Zhavoronkov, PhD (aka Aleksandrs Zavoronkovs) (43.8k followers) 1 h ago · 3 likes · 372 views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: arXiv.org.
More AI work like this
- "JEPA-Anything: Learning Predictive Models across Different Worlds" — @askalphaxiv
- A while ago, I realized that I needed to actually learn AI fundamentals, not just… — @0xdoug
- Impressive paper showing the impact of a good harness. — @dair_ai
- SoupFold from KAIST: no single co-folding model wins everywhere, so learn simple… — @biogerontology
- We built AI that actually knows aging biology. Longevity-LLM (9B params) beats OpenAI… — @biogerontology
- We need more examples like this in the open-source RL ecosystem — @adithya_s_k
- "Verbalizing Subliminal Learning Effects Using Text Optimization" — @askalphaxiv
- “What Does Privileged Information Add to On-Policy Self-Distillation?” — @askalphaxiv
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.7k posts from 4.9k X accounts over the last 14 days, 1.5k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 17:01 UTC. Full method.