我使用 Qwen3.5-4B 底座配 rank-8 LoRA 跑 RLCD,3.2 万道题、3 个种子,9 个 GPU 小时,花了约 $42。
This is a AI post classified by Jev as AI infra & evals (a model release), kept by the AI Radar because it carries real work, not commentary.
如何自己训练一个 Jev 一样的判断模型? 我使用 Qwen3.5-4B 底座配 rank-8 LoRA 跑 RLCD,3.2 万道题、3 个种子,9 个 GPU 小时,花了约 $42。 第三方 JevBench 120 道题,和 Jev 逐题打平。 更早用 Gemma E4B 训练的一轮,81.2 对 Jev 76.9,在中文上的判断准确率超过了 Jev。
Posted by AstroHan (1.3k followers) 1 h ago · 3 likes · 188 views · view the original post on X. Kept by the AI Radar as AI infra & evals.
More AI work like this
- Qwen-Image-2.1 on one Spark: day-0 Comfy, 23/23, H3 kept — @MichaelGannotti
- UPDATE we COOKING ! — @Tech2Wild
- just launched teams on http://classifier.dev! — @michael_chomsky
- we've been exploring how to use Jev to build a better harness — @hwchase17
- Step-5-preview inference speed has improved, time to run more evals. — @ItsmeAjayKV
- 95 tok/s 🤯 — @MiaAI_lab
- I BUILT A ROUTING DESK FOR 188 AI MODELS TO SEE HOW MUCH OF AN 11-TOOL, $277/MO STACK I… — @gippp69
- 6 GiB. That is the entire KV bill for a 262K context window on Qwen3.8-Flash-Next,… — @cozybearlog
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.8k posts from 4.9k X accounts over the last 14 days, 1.5k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 17:39 UTC. Full method.