AI Radar
Support
LiveUpdated 2026-09-23 14:57 UTC

It has been a really interesting experience to build an eval suite for System One Models…

It has been a really interesting experience to build an eval suite for System One Models like Jev. My first two…

This is a AI post classified by Jev as AI infra & evals (a tool drop), kept by the AI Radar because it carries real work, not commentary.

It has been a really interesting experience to build an eval suite for System One Models like Jev. My first two attempts didn't work out. But I learned a lot in the process, and since Jev is so fast, and cheap, it's a lot easier to run quick tests than with an LLM. I'm on attempt number three with this eval suite, and thought I'd share more about it. The eval suite is called @VulcanBench Verdict, and here's the high-level on how it works:

Posted by Morgan (46.1k followers) 1 h ago · 14 likes · 923 views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.4k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 14:57 UTC. Full method.