NVIDIA brings its entire speech stack local: ASR, TTS and codec now run on-device
NVIDIA has released speech synthesis, recognition and compression models in GGUF format, runnable locally via NeMo-Speech.cpp. This move enables on-premise voice pipelines, cutting cloud dependency and bolstering audio data control.