NVIDIA research papers are on fire recently!
This is a AI post classified by Jev as Voice AI (research), kept by the AI Radar because it carries real work, not commentary.
NVIDIA research papers are on fire recently! Here is another interesting paper where they give full-duplex speech models tool calls. (bookmark it) Commercial duplex voice models complete 31 to 51 percent of grounded customer-service tasks under clean conditions. Text agents like GPT-5 reach 85 percent on the same tasks in text mode. Most of what a voice agent loses, it loses in the speech pipeline. The fix routes the decision out of the speech model. The duplex frontend learns to emit a delegation token, forwards streaming transcripts to a text backend LLM for the tool call, and receive
Posted by elvis (320.4k followers) 1 h ago · 10 likes · 2k views · view the original post on X. Kept by the AI Radar as Voice AI. Tools mentioned: academy.dair.ai.
More AI work like this
- NetEase-Youdao releases Confucius4-R2T2, a true-streaming ASR model for live captions,… — @ModelScope2022
- New model: qwen-audio-3.1-realtime-plus — @aitrackerbot
- Meet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation… — @alibaba_cloud
- SpaceXAI just released Grok Voice Transcribe 2.0.....and it already ranks #1 for BOTH… — @XFreeze
- 开源社区又在整大活: — @mylifcc
- なんてこった!! — @studio_yebisu
- This is the best open source alternative to ElevenLabs and WisprFlow that you can runs… — @hasantoxr
- Meet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation… — @Alibaba_Qwen
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 34.2k posts from 4.8k X accounts over the last 14 days, 1.4k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-20 10:08 UTC. Full method.