Browser Use Bench v2 Pareto frontier got completely redrawn today
This is a AI post classified by Jev as AI agents (a model release), kept by the AI Radar because it carries real work, not commentary.
Browser Use Bench v2 Pareto frontier got completely redrawn today > Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5x cheaper) > GPT‑6 Luna xhigh: 57.6 (22x cheaper than Opus) OpenAI is in its own league 🔥 All models available to try on our cloud. [x-axis is log scale]
Posted by Browser Use (47.6k followers) 1 h ago · 6 likes · 1k views · view the original post on X. Kept by the AI Radar as AI agents.
More AI work like this
- We just wrapped up our first ever Ship It conference for healthcare software engineers. — @nikillinit
- Meta Closed Muse vs Open Muse. — @0x0SojalSec
- Breaking: Luna and Sol are INSANE for Browser Use 👀 — @gregpr07
- Tesla Full Self-Driving with Jev. — @0x0SojalSec
- Please join me Thursday to walk through a live webinar: Building a Credit AI Workflow. — @FundamentEdge
- Grok Bot just got a major upgrade over the last few days — @XFreeze
- Some recent quality-of-life improvements to Grok Bot. — @bot
- I made my own visualizer for the major models we're using for agents today. — @theo
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 43.6k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 01:02 UTC. Full method.