Scientific discovery poses significant challenges for measuring an AI system, as…
This is a AI post classified by Jev as AI research (a model release), kept by the AI Radar because it carries real work, not commentary.
Scientific discovery poses significant challenges for measuring an AI system, as effective signals on discovery outcome are slow, expensive and scarce. We present TRACES, a benchmark and a methodology to evaluate discoverative AI systems based on their process, along six fundamental capabilities: T – Tool usage R – Recovery from errors A – Alternatives management C – Coherence over many steps E – Evidence loyalty S – Scope descriptions that are clear We evaluated 4 frontier open-sourced models: @Kimi_Moonshot kimi-k3, @Alibaba_Qwen qwen3.8-max, @deepseek_ai deepseek-v4.1-flash, and
Posted by Apodex (5.5k followers) 1 h ago · 2 likes · 189 views · view the original post on X. Kept by the AI Radar as AI research. Tools mentioned: The TRACES Leaderboard.
More AI work like this
- The absolutely beautiful free hands-on book "Physics-Based Deep Learning Book" by… — @burkov
- Every mature systems project eventually needs a teaching version. — @PyTorch
- update: it broke :( — @pentestduck
- a model trained for 24 days has now resolved more than 100 long-standing open problems… — @IterIntellectus
- SpecOpt (Nguyen & Ji, UIUC) reframes selectivity as its own design task: dock vs target… — @biogerontology
- 🚨wtf. OpenAI says its new internal model it began training just 24 days ago has already… — @choblin29
- STEM educators, this day is designed for you - https://bit.ly/4gwaWmH — @argonne
- Found a great production use case for Jev. — @omarsar0
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.3k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 21:52 UTC. Full method.