Since we launched the first decision model in July, we've been concerned that models…
This is a AI post classified by Jev as Frontier models (a launch), kept by the AI Radar because it carries real work, not commentary.
Since we launched the first decision model in July, we've been concerned that models like ours (and now Jev) can't answer basic logic that a 12yo can solve. A ball is under cup A. Swap cups A and B. Swap cups B and C. Swap cups A and B. Swap cups A and C. Which cup is the ball under? Do you really care about a 200ms answer if it gets basic stuff confidently wrong? Today @levantolabs releases Sage1: – It's multimodal: it can see! – It reasons, but only when needed. So it knows where the cup is :) But there are downsides.
Posted by Marco De Rossi (6.6k followers) 1 h ago · 47 likes · 1.8k views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- Tested no-CoT abilities of GPT-5.6 Luna/Sol vs GPT-6 Luna/Sol - the GPT-6 series shows a… — @N8Programs
- Anthropic lança Claude Opus 5.5 com desempenho do Fable 5.1, mas incluído na assinatura… — @Tec_Mundo
- Grok 4.7 is insanely fast. It turned a simple prompt into this fun mini-game in just… — @cb_doge
- MiniMax tokenizer, M3.1? 👀 — @eliebakouch
- We asked Claude Opus 5.5, GPT-6 Sol, Sol Pro, GPT-6 Luna and Luna Pro for the single… — @eachlabs
- OpenAI GPT‑6 Sol and Luna are on the APEX leaderboards. — @mercor
- Gemini 3.8 Flash tops this chess benchmark — @HarshithLucky3
- GPT-6 Sol isn't just cheaper than GPT-5.6, it's also faster! — @AdamHoltererer
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.8k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 18:46 UTC. Full method.