Paige Bailey from Google DeepMind showcased the latest multimodal AI models like Gemini 3.1 and Genie 3, demonstrating their capabilities in text, image, audio, video, and code processing through the AI Studio platform for building and deploying versatile AI-powered applications. She highlighted features such as real-time interaction with Gemini Live, dynamic virtual world generation, and seamless integration with tools like Google Search, emphasizing accessibility, practical use cases, and future development plans to empower developers.
Paige Bailey from Google DeepMind presented an engaging and interactive session focused on building and deploying AI-powered applications using DeepMind’s latest models and tools. She began by introducing herself and highlighting the rapid release of several advanced AI models, including Gemini 3.1 (Flash, Pro, and Flashlight versions), Nano Banana 2 for image generation, LIIA 3 for music generation, Genie 3 for dynamic world building, and VO3.1 Light for video generation. These models support multimodal inputs and outputs, such as text, images, audio, video, and code, enabling versatile AI applications. Paige emphasized the unique capabilities of Gemini models, especially their ability to handle multiple modalities simultaneously and their integration with tools like Google Search and Maps for grounded, up-to-date information.
Paige demonstrated AI Studio, DeepMind’s platform for experimenting with these models, which is accessible via personal Gmail accounts and supports various configurations like structured outputs, code execution, and URL context for enhanced grounding. She showcased practical examples, such as analyzing YouTube videos to identify dinosaurs with timestamps and fun facts, comparing model performances using image analysis with bounding boxes, and generating SVG representations of images. The platform also supports code execution in a sandboxed Python environment, allowing models to perform complex data science tasks safely and efficiently.
One of the standout features Paige highlighted was Gemini Live, a real-time interactive model that can process screen sharing, video feeds, and audio to engage in dynamic conversations. This model supports multiple languages, accents, and dialects, and can perform tasks like describing screen content, answering questions, and even generating poetry in specified styles or accents. Gemini Live’s capabilities extend to robotics and augmented reality, where it can assist with object detection, navigation, transcription, and real-time interaction, showcasing its potential for diverse real-world applications.
Paige also introduced the Genie 3 model, designed for generating and navigating entirely new virtual worlds dynamically, pixel by pixel, without traditional game engines. She demonstrated creating a custom world with unique characters and environments, emphasizing the model’s compositional nature and its ability to produce immersive experiences. Although Genie 3 is not yet available as an API, it can be accessed through an ultra subscription in select countries. Additionally, Paige showed how AI Studio supports building and deploying custom apps with features like database integration and authentication, exemplified by a bookshelf cataloging app that uses image recognition and Google Search grounding to identify and store book details.
Throughout the session, Paige underscored the accessibility and versatility of DeepMind’s AI ecosystem, encouraging experimentation with free-tier models and APIs. She highlighted the seamless integration of generative media models for music, video, and image generation, and the availability of comprehensive code snippets for replicating experiments in various programming languages. The presentation concluded with a Q&A, where Paige addressed questions about future developments, including plans for an AI Studio app and integrations with tools like OpenClaw, reinforcing DeepMind’s commitment to empowering developers with cutting-edge, multimodal AI technologies.