$NVDA OPEN-SOURCES NEMOTRON 3, LETTING VOICE AGENTS TRACK WHO’S SPEAKING
This is a AI post classified by Jev as Voice AI (a model release), kept by the AI Radar because it carries real work, not commentary.
$NVDA OPEN-SOURCES NEMOTRON 3, LETTING VOICE AGENTS TRACK WHO’S SPEAKING Nvidia released Nemotron 3 Diarization, a 100M-parameter open-weight model designed to track who is speaking, when they speak, and when multiple people talk at once in real time. This is really improtant since speech-to-text can only tell an AI what was said, but without diarization it can struggle to know which person made a request, commitment or interruption. Nemotron 3 can track up to 8 speakers, versus 4 for Nvidia’s previous streaming model, while keeping speaker labels consistent across a conversation. It also
Posted by Wall St Engine (194.3k followers) 1 h ago · 53 likes · 14.7k views · view the original post on X. Kept by the AI Radar as Voice AI.
More AI work like this
- 🔈 Holy! You need to listen to this, and then run through SynthID or something to… — @altryne
- ElevenLabs is coming to @Muse. — @ElevenLabs
- Change the language, keep your voice. — @cartesia
- ChatGPT Voice mode is now supporting Plugins and can be used in ChatGPT Work! — @testingcatalog
- Gemini API skill has been updated courtesy of @_philschmid 🫶 — @thorwebdev
- voice is undefeated for studying — @daradoescode
- ...and it even works in other languages! 🇩🇪 — @DynamicWebPaige
- Gemini 3.8 Flash TTS and Flash-Lite TTS from Google (@Google) are now available on Merge… — @merge_api
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 45.1k posts from 5k X accounts over the last 14 days, 1.9k tools, 19 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 20:32 UTC. Full method.