AI Radar
Support
LiveUpdated 2026-09-21 21:52 UTC

Scientific discovery poses significant challenges for measuring an AI system, as…

Scientific discovery poses significant challenges for measuring an AI system, as effective signals on discovery…

This is a AI post classified by Jev as AI research (a model release), kept by the AI Radar because it carries real work, not commentary.

Scientific discovery poses significant challenges for measuring an AI system, as effective signals on discovery outcome are slow, expensive and scarce. We present TRACES, a benchmark and a methodology to evaluate discoverative AI systems based on their process, along six fundamental capabilities: T – Tool usage R – Recovery from errors A – Alternatives management C – Coherence over many steps E – Evidence loyalty S – Scope descriptions that are clear We evaluated 4 frontier open-sourced models: @Kimi_Moonshot kimi-k3, @Alibaba_Qwen qwen3.8-max, @deepseek_ai deepseek-v4.1-flash, and

Posted by Apodex (5.5k followers) 1 h ago · 2 likes · 189 views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: The TRACES Leaderboard.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.3k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 21:52 UTC. Full method.