Reminder: Claude Code can test if a skill or plugin actually improves Claude's answers.
This is a AI post classified by Jev as AI coding agents (a tutorial), kept by the AI Radar because it carries real work, not commentary.
Reminder: Claude Code can test if a skill or plugin actually improves Claude's answers. > claude plugin eval init Tell Claude what a good result looks like in your plugin's folder and it writes the test cases for you. > run claude plugin eval .
Posted by Addy Osmani (422.4k followers) 1 h ago · 227 likes · 12.9k views · view the original post on X. Kept by the AI Radar as AI coding agents.
More AI work like this
- ZCode, the official harness for GLM, is now OPEN SOURCE 🔥 — @Hesamation
- @Da7_Tech — @Da7_Tech
- 这太离谱了: — @AstroHanRay
- Amazing result, alas this score is probably not super informative — @teortaxesTex
- Looks like the MiMo RL runs have completed. The pro model went from 58.41 to 72.57 on… — @nrehiew_
- For all the CLI users: Update Codex and turn on Voice mode in /experimental. — @derrickcchoi
- Grok Build v1.0.40 is out with a few practical updates. The main item is miscellaneous… — @blankspeaker
- I just spent 30 mins with Astra working on 1 issue. — @uzairansar
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 38.1k posts from 4.9k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 08:02 UTC. Full method.