AI Radar
Support
LiveUpdated 2026-09-19 18:08 UTC

Jev agent got 1/20 vs 17/20 for BrowserCode + Luna on our long horizon task benchmarks.

Jev agent got 1/20 vs 17/20 for BrowserCode + Luna on our long horizon task benchmarks. The speed is INSANE. Feels…

This is a AI post classified by Jev as AI agents (a tool drop), kept by the AI Radar because it carries real work, not commentary.

Jev agent got 1/20 vs 17/20 for BrowserCode + Luna on our long horizon task benchmarks. The speed is INSANE. Feels like early Browser Use.(loads of potential, unsolved problems) Browser use is very complex state space search. Very often you just have to think hard or go back. A model with 0 reasoning ability simply can't do that (yet?). Can we get that behavior with better memory + search, without adding a reasoning model? Really hope I can make this work.

Posted by Gregor Zunic (30k followers) 13 h ago · 111 likes · 14.6k views · view the original post on X. Kept by the AI Radar as AI agents.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.