Voice agents with Realtime Video — Sidney Primas, LemonSlice

Sidney Primas, CTO of Lemon Slice, presents their mission to create photorealistic, real-time avatars that pass the Avatar Turing test by leveraging advanced video diffusion models focused on human expressions and emotions, enabling seamless integration with voice agents for applications like language learning and AI sales. The company emphasizes hardware optimization and innovative techniques to maintain continuous, high-quality avatar generation, with future plans to develop emotionally intelligent models and establish industry benchmarks through their own Avatar Turing test.

Sidney Primas, CTO and founder of Lemon Slice, presents the company’s mission to create photorealistic avatars that can pass the Avatar Turing test—making them indistinguishable from humans in real-time video calls. Lemon Slice recently partnered with Microsoft to bring a virtual Teddy Roosevelt avatar to life in a presidential library, showcasing full-body avatars capable of real-time interaction, including hand movements and object physics. This visual layer aims to enhance human-AI interaction by leveraging humans’ natural ability to process visual information alongside voice.

Lemon Slice’s approach differs from other avatar companies by focusing on training world models specifically centered on humans. Although this method is more challenging to develop and deploy, it offers emergent capabilities such as full-body movement, micro-expressions, and realistic emotional responses. Their product can generate avatars from a single image, supporting various styles from photorealistic to cartoonish, and integrates seamlessly with existing voice agents via an API, enabling applications like language learning and AI sales calls.

Technically, Lemon Slice builds its avatars using a custom-trained video diffusion model that heavily emphasizes audio input to capture emotions and facial expressions accurately. To achieve real-time interactivity, the model is designed to only consider past inputs, avoiding future data during video generation. They also drastically reduce the typical multi-step denoising process to a single step, enabling fast video synthesis. A significant challenge they overcame is error accumulation over long video sequences, which they address with a novel, undisclosed method, allowing continuous avatar generation for hours without noticeable degradation.

The company also highlights the importance of hardware optimization and “model hardness”—the complex orchestration of GPU and CPU tasks to maintain smooth, real-time video streaming without stutters or interruptions. This engineering feat allows Lemon Slice to offer video avatar generation at a cost comparable to voice models, making it viable for consumer applications. Currently, they are working on an emotional engine to improve avatars’ emotional awareness and responsiveness, aiming to produce more natural and contextually appropriate reactions during conversations.

Looking ahead, Sidney envisions a future where a single end-to-end EQ (emotional intelligence) model processes user audio and video to generate avatar responses with high emotional fidelity, while a separate IQ (intelligence) model handles complex reasoning and tool use. This layered approach could become mainstream within two to three years. Lemon Slice plans to develop and publish its own Avatar Turing test to benchmark progress. The company is actively hiring and invites those interested in tackling these technical challenges to join their team.