AI Radar
Support
LiveUpdated 2026-09-19 18:45 UTC

Bonsai 2 has been evaluated with a low thinking budget for xhigh.

Bonsai 2 has been evaluated with a low thinking budget for xhigh. Quantization errors really show their impact on long…

This is a AI post classified by Jev as Frontier models (a model release), kept by the AI Radar because it carries real work, not commentary.

Bonsai 2 has been evaluated with a low thinking budget for xhigh. Quantization errors really show their impact on long sequences, and Qwen3.8 27B often needs more than 81K tokens to complete its answer. For coding problems, like in LiveCodeBench, this is not enough. Expect some surprises for long-horizon agentic tasks. It's probably not as good as the model card says. Remarkable work nonetheless, as always.

Posted by Benjamin Marie (6.8k followers) 20 h ago · 51 likes · 3.5k views · view the original post on X. Kept by the AI Radar as Frontier models.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 27.2k posts from 4.8k X accounts over the last 14 days, 1.2k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:45 UTC. Full method.