Gus Martins and Ian Ballantyne from Google DeepMind present the Gemini family’s Gemma 4 series—open, efficient AI models designed for user ownership, customization, and deployment on diverse hardware from mobile devices to enterprise servers, enabling privacy, offline use, and cost-effective AI applications. They highlight the strategic importance of AI sovereignty, demonstrating practical benefits in areas like multimodal processing, agentic tasks, and industry-specific solutions, while encouraging community collaboration to enhance these accessible, adaptable models.
In this presentation, Gus Martins and Ian Ballantyne from Google DeepMind introduce the Gemini family of models, focusing on the recently released Gemma 4 series. Gus explains that while Gemini models represent Google’s most advanced AI hosted on their servers, Gemma models are designed as open models that users can own, customize, and run on their own hardware. This ownership is crucial for scenarios requiring data privacy, customization, or offline capabilities. The Gemma 4 lineup includes models optimized for mobile and IoT devices (E2B and E4B) as well as larger models (26B and 31B parameters) that balance performance with hardware accessibility, enabling powerful AI applications on more modest setups.
Gus highlights the efficiency of these models, noting that despite their relatively smaller size compared to some competitors, they deliver impressive intelligence per parameter. The models support multimodal inputs like text, vision, and audio, and can perform complex tasks such as coding, function calling, and reasoning. He emphasizes that while these models may not be the absolute most intelligent available, they are highly capable for many practical applications, especially where cost and hardware constraints are considerations. The models are accessible for free trial on platforms like AI.dev, encouraging experimentation and adoption.
Ian Ballantyne expands on the practical benefits of owning open models, particularly in the context of agentic AI tasks that generate high token usage, such as programming and document analysis. He discusses how running models locally on devices like laptops or mobile phones can reduce reliance on cloud services and token-based costs, shifting the expense to energy and hardware utilization. Ian demonstrates how Gemma 4 models can operate on mobile devices, enabling interactive AI that can control phone functions and process multimodal inputs in real-time, thus unlocking new use cases for edge computing and offline AI.
The presentation also covers enterprise applications, where owning and fine-tuning models like Gemma can provide tailored solutions for specific industries, such as healthcare with the MedGemma variant. Ian showcases a demo running the 26B model on a Mac with 48GB of unified memory, performing multilingual translations locally. He stresses the importance of evaluating models based on specific tasks and workflows, encouraging users to integrate Gemma models into existing AI pipelines using OpenAI-compatible interfaces. This approach allows organizations to balance performance, cost, and control according to their unique needs.
In conclusion, Gus and Ian emphasize the strategic value of AI sovereignty—owning and controlling AI models to ensure privacy, customization, and resilience against service disruptions. They advocate for experimentation with Gemma models across devices and scales, from mobile phones to enterprise servers, highlighting the flexibility and efficiency these models offer. The speakers invite feedback and collaboration from the community to refine and expand the capabilities of open models, aiming to empower users with accessible, powerful AI tools that can be deployed and adapted independently.