Voice AI: when is the "Her" moment? — Neil Zeghidour, Gradium AI

Neil Zeghidour from Gradium AI highlighted the significant advancements in voice AI, including realistic voice cloning and the development of full-duplex speech-to-speech models like Moshi, which enable more natural, seamless conversations akin to those in the movie “Her.” However, he emphasized ongoing challenges such as latency, conversational fluidity, scalability, and privacy, advocating for integrated solutions that combine robust cascaded architectures with efficient, on-device technologies to achieve truly human-like voice interactions.

Neil Zeghidour from Gradium AI delivered a comprehensive talk on the current state and future challenges of voice AI, framed around the iconic “Her” movie analogy. Gradium AI focuses on developing foundational voice AI technologies such as speech-to-text, text-to-speech, speech-to-speech, and voice cloning, aiming to provide versatile building blocks for voice applications rather than vertical-specific solutions. Neil highlighted the impressive progress in voice cloning, demonstrating how just 10 seconds of recorded speech can be used to generate highly realistic synthetic voices, showcasing the rapid advancements in the field.

Despite these advancements, Neil emphasized that the voice AI experience today still falls short of the seamless, natural conversations depicted in “Her.” Current systems often rely on cascaded architectures—speech-to-text, large language models (LLMs), and text-to-speech—which introduce latency and limit conversational fluidity. He pointed out that while text-based intelligence has improved, the integration with voice remains challenged by delays, especially when external tool calls are involved, which can take several seconds and disrupt the natural flow of dialogue.

Neil also discussed the limitations of speech-to-speech models, which aim to reduce latency by directly converting spoken input to spoken output. However, most existing models operate in half-duplex mode, meaning they cannot handle overlapping speech or natural conversational interruptions, which are essential for human-like interaction. Gradium AI’s Moshi model, a full-duplex speech-to-speech system, was presented as a breakthrough that allows simultaneous listening and speaking, enabling more natural and robust conversations even in noisy or multi-speaker environments.

Another critical challenge Neil addressed is scalability and cost. Voice AI remains expensive to deploy at scale, particularly due to the computational demands of text-to-speech synthesis. To tackle this, Gradium AI developed Gradium Phonon, an on-device TTS model optimized to run efficiently on smartphone CPUs, offering high-quality voice synthesis without relying on costly cloud APIs. This approach not only reduces operational costs but also enhances user privacy by keeping data processing local, addressing growing concerns about data security in voice applications.

In conclusion, Neil argued that while voice AI has made significant strides, the journey to achieving the natural, intelligent, and reliable voice interactions portrayed in “Her” is ongoing and complex. The future lies in combining the robustness and intelligence of cascaded systems with the fluidity and immediacy of full-duplex speech-to-speech models, alongside scalable, privacy-conscious deployment strategies. Gradium AI invites collaboration and innovation to push the boundaries of voice AI and bring the vision of truly human-like voice assistants closer to reality.