AI Radar
Support
LiveUpdated 2026-09-22 20:50 UTC

StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming…

StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming Speech to Text with 1.7% WER, a…

This is a AI post classified by Jev as Voice AI (a model release), kept by the AI Radar because it carries real work, not commentary.

StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming Speech to Text with 1.7% WER, a notable improvement on StepAudio 2.5 ASR (4.7%) StepAudio 3 ASR is StepFun's new Speech to Text model, available through the StepFun API for non-streaming transcription, and the first StepFun model to reach the top of our non-streaming Speech to Text leaderboard. Key takeaways ➤ Accuracy: StepAudio 3 ASR scores 1.7% on the AA-WER Index, effectively tied with Alibaba's Fun-Realtime-ASR-preview at 1.7%. It leads on AA-AgentTalk, our held-out voice-agent dataset, at 1.4%, but

Posted by Artificial Analysis (148.6k followers) 1 h ago · 56 likes · 6.5k views · view the original post on X. Kept by the AI Radar as Voice AI.

More AI work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 43.1k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 20:50 UTC. Full method.