Why can't ChatGPT Voice set a timer? | Voice AI expert explains

The video explains that ChatGPT’s voice capabilities struggle with tasks like setting timers due to limitations in autoregressive speech models that lack precise timing mechanisms and do not effectively utilize token duration data, leading to inaccuracies and inconsistent responses. Voice AI expert Jaman Gao highlights challenges in reliability, ethical concerns, and the distinction between general-purpose and specialized AI solutions, while expressing optimism about future advancements like diffusion models and the importance of building robust, customizable voice AI systems.

The video explores the limitations and quirks of ChatGPT’s voice capabilities, particularly focusing on why it struggles with tasks like setting timers accurately or maintaining conversational flow, such as wanting the last word. Voice AI expert Jaman Gao explains that many speech-to-speech models are designed primarily as voice assistants that prioritize listening over interrupting, which affects their responsiveness and timing. The models operate autoregressively, generating tokens step-by-step without a precise sense of timing, which complicates tasks like beatboxing or keeping rhythm. While music generation models tend to use hierarchical structures for timing, voice models lack this, leading to inconsistencies.

The discussion highlights that although the technical data for timing (like token duration) exists within the system, current language models do not effectively utilize this information to measure or track time accurately. This results in errors when users ask the AI to start and stop timers or measure durations, often producing imaginative or incorrect results. The expert suggests that integrating tool calls—specialized functions that handle tasks like timing—could improve performance, but this has not been fully implemented yet. Experiments with newer models like GPT Live show some improvement but still fall short of perfect accuracy.

The video also touches on the challenges of voice AI in handling complex or sensitive interactions, such as emergencies or emotional conversations. For example, when users simulate being in danger (like sinking in quicksand), the AI often treats these scenarios as jokes rather than emergencies, raising concerns about reliability and ethical implications. The expert notes that while voice AI can help address loneliness or elderly care, the unpredictability and lack of control over AI responses pose risks, especially for vulnerable users. The balance between creating engaging demos and building robust, trustworthy systems remains a significant challenge.

Another key point is the difference between horizontal AI companies, like OpenAI providing general-purpose models, and vertical companies that tailor AI solutions for specific industries or use cases. Horizontal models often have broad capabilities but can produce errors or inappropriate responses, while vertical companies face the difficult task of ensuring reliability and safety in specialized applications like therapy or emergency response. The expert suggests that many companies rely on combining APIs from different providers to build complete products, which can complicate robustness and increase costs but also allows for enterprise-level customization.

Finally, the video concludes with reflections on the future of voice AI technology. The expert expresses excitement about emerging architectures like diffusion models that can generate many audio tokens simultaneously, potentially improving efficiency and quality. However, challenges remain in standardizing interfaces across different speech recognition providers and creating seamless, reliable voice assistants. Despite current limitations, the expert remains optimistic about the potential for innovation and encourages developers to build more robust, debuggable voice AI systems using available tools and open-source resources.