How far are we from "Her"?

The video discusses the advancements in AI voice assistants, highlighting the shift from traditional cascade models to innovative end-to-end, full duplex systems like OpenAI’s GPT Live and Moshi that enable more natural, simultaneous conversational interactions akin to the AI “Samantha” from Her. While significant progress has been made, practical deployment challenges and ethical considerations mean that truly human-like AI companionship remains a complex and ongoing pursuit.

The video explores how close current AI voice assistants are to the fictional AI “Samantha” from the 2013 movie Her, focusing on recent advancements in voice AI technology. It contrasts traditional voice assistants, which operate in a cascade model—processing speech in stages from voice detection to transcription to language modeling and finally speech synthesis—with newer end-to-end models that handle audio input and output more holistically. The cascade approach, while practical and widely used, suffers from latency and loss of acoustic nuance, whereas end-to-end models like OpenAI’s GPT40 can process multiple modalities simultaneously, reducing delay and improving naturalness.

A major breakthrough discussed is the development of full duplex voice assistants, which can listen and speak simultaneously, unlike earlier half-duplex systems that alternate between listening and speaking. This capability allows for more natural conversational dynamics, such as interruptions and backchanneling, where the AI can acknowledge the user without waiting for its turn. OpenAI’s GPT Live and the open-source Moshi assistant from the Paris-based nonprofit QAI exemplify this new class of interaction models, which treat conversations as overlapping audio streams rather than discrete turns, enabling more fluid and human-like interactions.

Moshi’s architecture is highlighted for its innovative approach to audio tokenization, breaking down speech into semantic and acoustic tokens that are processed in a hierarchical manner. This design allows Moshi to perform multiple tasks—speech-to-text, text-to-speech, and conversational interaction—within a single unified model. Additionally, Moshi Rag extends this by integrating asynchronous retrieval of external information during conversations, maintaining a natural flow even when the AI needs to access background knowledge, a feature crucial for practical applications like customer support.

Despite the promise of full duplex, end-to-end models, the video explains why cascade systems remain dominant in the industry. Cascades allow companies to leverage existing large language models without retraining entire pipelines, making them more practical and easier to deploy in the short term. However, leading AI labs are investing in end-to-end models for their long-term potential to deliver more seamless and human-like voice assistants. The transition to these models will require advances in synthetic data generation and new evaluation metrics to measure conversational naturalness effectively.

In conclusion, while we are making significant strides toward AI assistants resembling Samantha from Her, the journey is ongoing and complex. Current AI systems excel in some areas but still fall short in others, reflecting a different kind of intelligence than humans possess. The video ends with a thoughtful reflection on the ethical and emotional implications of pursuing AI companionship, emphasizing that while AI can enhance services like customer support or therapy, genuine human connection and love remain irreplaceable by machines.