i've been trying to run warden security benchmarks against 4.7 since yesterday and it…
This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.
i've been trying to run warden security benchmarks against 4.7 since yesterday and it hasn't been great. the last grok model that was absolutely crushing it in vulnerability scanning is 4.5 and we still use it as base driver. both 4.6 and 4.7 disappointed, were spinning for two long, missing or "disproving" some real results and costing way too much. original numbers are here https://warden.sentry.dev/benchmarking (although we didn't add 4.7 yet).
Posted by Greg Pstrucha (1.2k followers) 55 min ago · 0 likes · 6 views · view the original post on X. Kept by the AI Radar as Frontier models.
More AI work like this
- GPT-6 Sol recreates The Nihilist Penguin in Blender — @aimlapi
- My prompt was "Generate Teenage Engineering's OP-1". — @zachdive
- Claude Opus 5.5 vs GPT-6 Sol in building game worlds. — @higgsfield_ai
- Opus 5.5 is now the default model in Adam. — @adamdotnew
- Opus 5.5 is the most fun model I’ve ever used. — @AiBreakfast
- BREAKING: OpenAI's GPT-6 Astra successfully drove a real Toyota Corolla through an… — @exec_sum
- BREAKING: OpenAI's GPT-6 Astra successfully drove a real Toyota Corolla through an… — @exec_sum
- Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture… — @EpochAIResearch
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 44.7k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 17:53 UTC. Full method.