AI Radar
Support
LiveUpdated 2026-09-21 18:11 UTC

Okay, finally finished my Muse Spark 1.3 benchmark.

Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done…Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done at @VulcanBench and it definitely looks like Muse has some serious issues at lower effort levels. I was able to run Astra through my v4 eval suite across every effort level in ~12 hours, Muse Spark had single tasks at low effort levels that hit the 10 hour mark, on just a single task. It took a total of two weeks for me to run the entire benchmark. Overall my assessment is that Muse is a promising model, but they'll have to figure out why it has issues at lower effort level

Posted by Morgan (46.1k followers) 2 h ago · 57 likes · 4.6k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 39.6k posts from 5k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 18:11 UTC. Full method.