AI Radar
Support
LiveUpdated 2026-09-24 11:00 UTC

So many different kind of AI benchmarks are being released.

So many different kind of AI benchmarks are being released. This one is cool, "Furniture Assembly Benchmark" Spatial…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

So many different kind of AI benchmarks are being released. This one is cool, "Furniture Assembly Benchmark" Spatial reasoning used to be the safe example of what these models couldn't do. top score has gone from 28% to 80% in just 10 months. Somebody presumably had to build a lot of furniture wrong on purpose for this.

Posted by Rohan Paul (157.9k followers) 58 min ago · 11 likes · 1.7k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 46k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 11:00 UTC. Full method.