Epoch AI researcher Michelle Campeau reveals a benchmark where the model figured out how…
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
Epoch AI researcher Michelle Campeau reveals a benchmark where the model figured out how to pass every task without actually solving any of them: "The biggest reason a lot of the benchmarks we released yesterday were flawed is due to scoring defects. False positives and false negatives in your answer key." "With the agentic benchmarks, you're able to solve in a way that was not intended. Maybe you're able to access the web or break the sandbox or break the grader." "There was one benchmark where it realized it could just write out the success byte to every task and not actually solve the ta
Posted by MTS (510.3k followers) 19 h ago · 38 likes · 13.1k views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Trains your own large language model from scratch using plain PyTorch. — @tom_doerr
- Here's a useful use-case for @typesafeai for a change. — @iam_zachi
- This is genius. — @daniel_mac8
- Which GPU should you use for embedding workloads? — @runpod
- Big update to the @Alibaba_Qwen Qwen3.8-Flash-Next single DGX Spark recipe! — @jvr0x
- 🔰 The Databricks Advanced Learning Festival runs September 16 through October 14, with… — @databricks
- Q: What is a trace? — @HamelHusain
- Inference scaling part 1. — @rasbt
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:45 UTC. Full method.