So many different kind of AI benchmarks are being released.
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
So many different kind of AI benchmarks are being released. This one is cool, "Furniture Assembly Benchmark" Spatial reasoning used to be the safe example of what these models couldn't do. top score has gone from 28% to 80% in just 10 months. Somebody presumably had to build a lot of furniture wrong on purpose for this.
Posted by Rohan Paul (157.9k followers) 58 min ago · 11 likes · 1.7k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- GPT 6 Sol beats Fable 5 on BridgeBench. At 1/5 of the cost. — @bridgemindai
- We just added GPT-6 Sol to MathArena! It comes in only just behind GPT-6 Astra, at… — @j_dekoninck
- Gemini 4 is coming sooner than expected — @notjazii
- 真忍不住了,都在说 Opus 5.5 很牛逼,特别是写作和3D能力 — @Saccc_c
- opus 5.5 made this gothic spaceship chess game and it looks majestic. — @Ananth7e
- Opus 5.5 medium is faster and cheaper than both Opus 4.6 and Fable 5.1 at high — @MaaSonder
- WHAT JEV IS AND HOW IT PROCESSES A SINGLE CALL — @DataChaz
- 🚨 Gemini 4 could actually come pretty soon then — @LuminaBench
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 46k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 11:00 UTC. Full method.