StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming…
This is a AI post classified by Jev as Voice AI (a model release), kept by the AI Radar because it carries real work, not commentary.
StepFun has released StepAudio 3 ASR, ranking #1 on the AA-WER Index for non-streaming Speech to Text with 1.7% WER, a notable improvement on StepAudio 2.5 ASR (4.7%) StepAudio 3 ASR is StepFun's new Speech to Text model, available through the StepFun API for non-streaming transcription, and the first StepFun model to reach the top of our non-streaming Speech to Text leaderboard. Key takeaways ➤ Accuracy: StepAudio 3 ASR scores 1.7% on the AA-WER Index, effectively tied with Alibaba's Fun-Realtime-ASR-preview at 1.7%. It leads on AA-AgentTalk, our held-out voice-agent dataset, at 1.4%, but
Posted by Artificial Analysis (148.6k followers) 1 h ago · 56 likes · 6.5k views · view the original post on X. Kept by the AI Radar as Voice AI.
More AI work like this
- Do you want early access to Muse phone agents that can make calls for you? I think I… — @WaitWhat
- Announcing Multilingual Text to Speech Arena Leaderboards, comparing leading TTS models… — @ArtificialAnlys
- Announcing Multilingual Text to Speech Arena Leaderboards, comparing leading TTS models… — @ArtificialAnlys
- This is mindblowing. Your Tesla can now manage your inbox, calendar, and tasks while you… — @VaibhavSisinty
- Use Opus 5.5 on MEDIUM! — @AdamHoltererer
- Gradium TTS beta model is out now to define what fast means. — @GradiumAI
- New model: gemini-3.8-flash-tts — @aitrackerbot
- New model: gemini-3.8-flash-lite-tts — @aitrackerbot
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 43.1k posts from 5k X accounts over the last 14 days, 1.8k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 20:50 UTC. Full method.