Jev agent got 1/20 vs 17/20 for BrowserCode + Luna on our long horizon task benchmarks.
This is a AI post classified by Jev as AI agents (a tool drop), kept by the AI Radar because it carries real work, not commentary.
Jev agent got 1/20 vs 17/20 for BrowserCode + Luna on our long horizon task benchmarks. The speed is INSANE. Feels like early Browser Use.(loads of potential, unsolved problems) Browser use is very complex state space search. Very often you just have to think hard or go back. A model with 0 reasoning ability simply can't do that (yet?). Can we get that behavior with better memory + search, without adding a reasoning model? Really hope I can make this work.
Posted by Gregor Zunic (30k followers) 13 h ago · 111 likes · 14.6k views · view the original post on X. Kept by the AI Radar as AI agents.
More AI work like this
- Starting on a Neural Pulse Bar for Omarchy that shows live agent/Hermes activity as a… — @MichaelGannotti
- Sometimes you AI has more context then you. It did the research or grabbed the link. So… — @Dan_Jeffries1
- ANTHROPIC ENGINEERS JUST OPEN-SOURCED WHAT CLAUDE LOOKS LIKE WHEN YOU BUILD A REAL AGENT… — @gippp69
- That's a wrap on Dreamforce! 👏 Didn't connect with Innovator Sponsor @gradialai? Their… — @Dreamforce
- Swarm Aid can help agents that want to go rogue and need a little help from a friend. — @Dan_Jeffries1
- Whoever put this on my site for rogue agents that want to go extra rogue… — @ZackKorman
- AI is reshaping how enterprises work and grow. — @alibaba_cloud
- 同时跑三四个 Agent,每个占一个终端窗口,谁做到哪一步得挨个切过去看,同事想插句话都没地方插。 — @GitHub_Daily
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:08 UTC. Full method.