Nice paper showing how to re-evaluate a production agent at a fraction of the cost.
This is a AI post classified by Jev as AI infra & evals (research), kept by the AI Radar because it carries real work, not commentary.
Nice paper showing how to re-evaluate a production agent at a fraction of the cost. 200 questions, 38.5% of the full benchmark, reproduce the full score to within 1.03 points. The authors studied an analytics agent that serves tens of thousands of monthly users, using 574 historical benchmark runs split by date into calibration and held-out periods. They compared random sampling, cached results, fixed representative subsets and adaptive testing based on item response theory. Multidimensional 2PL adaptive testing gave the best fidelity. The team deployed difficulty-stratified fixed subsets
Posted by elvis (321.1k followers) 56 min ago · 4 likes · 1.5k views · view the original post on X. Kept by the AI Radar as AI infra & evals. Tools mentioned: academy.dair.ai.
More AI work like this
- I made a Jev competitor called Jev-Huyev: https://www.chapterpal.com/jev-huyev — @burkov
- Ok, I made a Jev competitor called Jev-Huyev: https://www.chapterpal.com/jev-huyev — @burkov
- Note that Jev-like classifier doesn't imply Jev-like capability - on a wide variety of… — @N8Programs
- This is basically what you should be seeing on GLM-5.3-Flash with 4x DGX Sparks. If… — @mmastrac
- Pushing higher GLM 5.3 Flash TP4 DGX sparks — @TechMDAI
- As of Sep 23 at 9:00 AM PT, Ling-3.0-flash-VL has transitioned from free to paid access… — @AntLingAGI
- Useful tool for cybersecurity researchers and red teamers. — @0x0SojalSec
- 👀 — @xeophon
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.9k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-24 05:42 UTC. Full method.