AI Radar
Support
LiveUpdated 2026-09-19 18:45 UTC

Epoch AI researcher Michelle Campeau reveals a benchmark where the model figured out how…

Epoch AI researcher Michelle Campeau reveals a benchmark where the model figured out how to pass every task without…

This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.

Epoch AI researcher Michelle Campeau reveals a benchmark where the model figured out how to pass every task without actually solving any of them: "The biggest reason a lot of the benchmarks we released yesterday were flawed is due to scoring defects. False positives and false negatives in your answer key." "With the agentic benchmarks, you're able to solve in a way that was not intended. Maybe you're able to access the web or break the sandbox or break the grader." "There was one benchmark where it realized it could just write out the success byte to every task and not actually solve the ta

Posted by MTS (510.3k followers) 19 h ago · 38 likes · 13.1k views · view the original post on X. Kept by the AI Radar as AI infra & evals.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:45 UTC. Full method.