Okay, finally finished my Muse Spark 1.3 benchmark.

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done at @VulcanBench and it definitely looks like Muse has some serious issues at lower effort levels. I was able to run Astra through my v4 eval suite across every effort level in ~12 hours, Muse Spark had single tasks at low effort levels that hit the 10 hour mark, on just a single task. It took a total of two weeks for me to run the entire benchmark. Overall my assessment is that Muse is a promising model, but they'll have to figure out why it has issues at lower effort level
Posted by Morgan (46.1k followers) 2 h ago · 57 likes · 4.6k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- Sam Altman when he saw "Grok 4.7 is 83% cheaper than GPT-6 Astra" — @0x0SojalSec
- Grok 4.7 from xAI (@SpaceXAIx) is now available on Merge Gateway. — @merge_api
- interesting GPT-6 Luna leak in the Codex code — @haider1
- Grok 4.7 Is Actually Insane 🔥 — @Priyannkaaaa
- Got a question from @ALL_Jonah about @grok 4.7 and how people can access it now and why… — @MichaelGannotti
- $0.20 vs $2.05. — @ai_for_success
- we evaluated @grok-4.7 on an internal sample of 22 knowledge-work tasks across engg,… — @himanshustwts
- Before you meet MiMo-V2.6, step into its world. — @XiaomiMiMo
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 39.6k posts from 5k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 18:11 UTC. Full method.