It has been a really interesting experience to build an eval suite for System One Models…
This is a AI post classified by Jev as AI infra & evals (a tool drop), kept by the AI Radar because it carries real work, not commentary.
It has been a really interesting experience to build an eval suite for System One Models like Jev. My first two attempts didn't work out. But I learned a lot in the process, and since Jev is so fast, and cheap, it's a lot easier to run quick tests than with an LLM. I'm on attempt number three with this eval suite, and thought I'd share more about it. The eval suite is called @VulcanBench Verdict, and here's the high-level on how it works:
Posted by Morgan (46.1k followers) 1 h ago · 14 likes · 923 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- We want to sponsor Rust maintainers and open-source projects across web development,… — @encoredotdev
- Spent a bit of time tuning NCCL for TP=4 GLM 5.3 Flash on 4x DGX Sparks. Bumped to… — @mmastrac
- GLM 5.3 Flash EXL3 on 2x DGX Spark: a new cache update is merged! — @plotarmordev
- We got @DeepSeek-V4.1-Flash running in EXL3 on 4× Sparks! — @cfontes
- Qualcomm announces it most powerful Snapdragon 8 Elite Extreme Gen 6 chip built for… — @CNET
- The https://MLX.fast community has finally doubled the speed of Qwen 3.8 Flash Next to… — @TheDavidTai
- Have you ever had a local coding agent flying at the start of a session and crawling an… — @jvr0x
- 🚀 MCDMA is working between my Mac Studio & DGX Sparks, and now it's time to bring it to… — @ashxhart
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.4k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 14:57 UTC. Full method.