StepAudio 3 Gen: a unified audio generation model
This is a AI post classified by Jev as Voice AI (a model release), kept by the AI Radar because it carries real work, not commentary.
StepAudio 3 Gen: a unified audio generation model Supports TTS, voice design, vocals, sound effects, music, and mixtures — using a discrete autoregressive generator over RVQ tokens. Achieves state-of-the-art performance on both TTS and voice design.
Posted by DailyPapers (21.4k followers) 2 days ago · 30 likes · 2.4k views · view the original post on X. Kept by the AI Radar as Voice AI.
More AI work like this
- A little experiment into the compression of ASR models for some later ideas. The QAD… — @DJLougen
- Day after voice calls, Grok Bot can send voice notes. I talk live when I’m in it — then… — @MichaelGannotti
- NetEase-Youdao releases Confucius4-R2T2, a true-streaming ASR model for live captions,… — @ModelScope2022
- NVIDIA research papers are on fire recently! — @omarsar0
- New model: qwen-audio-3.1-realtime-plus — @aitrackerbot
- Meet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation… — @alibaba_cloud
- SpaceXAI just released Grok Voice Transcribe 2.0.....and it already ranks #1 for BOTH… — @XFreeze
- 开源社区又在整大活: — @mylifcc
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 38k posts from 4.9k X accounts over the last 14 days, 1.6k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-21 07:04 UTC. Full method.