1/ Introducing Cua-Bench-S1, a benchmark for decision models built for computer use,…
This is a AI post classified by Jev as AI agents (a model release), kept by the AI Radar because it carries real work, not commentary.
1/ Introducing Cua-Bench-S1, a benchmark for decision models built for computer use, including Jev. We're releasing the first generation of Cua-S1 models, with two checkpoints: Cua-S1-Nano-0.1 and Cua-S1-4B-0.1 Cua-S1-Nano-0.1: https://huggingface.co/cua-ai/cua-s1-nano-0.1 Cua-S1-4B-0.1: https://huggingface.co/cua-ai/cua-s1-4b-0.1 Repo: https://github.com/trycua/cua
Posted by Cua (19.2k followers) 1 h ago · 157 likes · 7k views · view the original post on X. Kept by the AI Radar as AI agents. Tools mentioned: cua-s1-nano-0.1, cua-s1-4b-0.1, cua.
More AI work like this
- looking to leverage Grok Build to fine tune and tweak my Harness by Autonomous — @MichaelGannotti
- Meta ❤️ Shopify — @testingcatalog
- I have my Muse using Grok 4.7 for difficult tasks and for checking Muses work. — @WaitWhat
- A drone needs power. — @arc
- The latest flashpoint in the B2A era is Amazon who just blocked agents like Muse and… — @stevehind
- Codos has launched a virtual Chief AI Officer, a software that can plan and run a… — @testingcatalog
- @openclaw — @openclaw
- NEW: Subagents can now point back to an earlier subagent's result instead of the parent… — @mastra
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 40.4k posts from 5k X accounts over the last 14 days, 1.7k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 23:00 UTC. Full method.