Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed discuss the limitations of current AI models, particularly their lack of genuine continual learning, and advocate for developing adaptive systems that learn continuously from real-world experience using advanced algorithms like continual backpropagation. They envision a future where AI, inspired by biological learning, can perform lifelong learning and abstraction across diverse domains, a vision their company Oak Lab aims to realize through innovative, efficient, and scalable approaches.

In this insightful conversation, Rich Sutton, a pioneer in reinforcement learning, and Khurram Javed discuss foundational ideas and future directions in AI research. Sutton emphasizes that continual learning is the natural and essential mode of intelligence, contrasting it with the current AI field’s fragmented approach that treats continual learning as a special case. Reflecting on his career, Sutton shares his conviction to focus on reinforcement learning even during the AI winter and personal health challenges, underscoring his belief that learning and goal-directed behavior are central to intelligence. He critiques the current AI paradigm that overly relies on static models like large language models (LLMs), which do not continue learning after training, and stresses the importance of systems that learn continually from experience.

The conversation delves into Sutton’s “bitter lesson,” which argues that human knowledge should not be overly injected into AI systems; instead, scalable learning methods that leverage computation are key to long-term progress. While acknowledging the breakthroughs of LLMs, Sutton and Javed caution against overreliance on synthetic data generation, highlighting the “big world hypothesis”—the idea that the real world is infinitely complex and cannot be fully captured by simulations or synthetic data. They argue that true continual learning requires agents to learn from their own real-world experiences rather than depending on human-curated data or simulations, which are limited by human expertise and simplifications.

Sutton and Javed also discuss the limitations of current AI systems, particularly the lack of genuine continual learning and the inability to update model weights dynamically during deployment. They explain that naive continual learning leads to catastrophic forgetting, but propose advanced algorithmic solutions like “continual backpropagation,” which involves metalearning step sizes and generating new features to enable sustained learning without losing prior knowledge. This approach, they believe, could revolutionize AI by allowing models to learn continuously and adaptively, overcoming the static nature of today’s LLMs and enabling personalized, efficient learning.

Drawing inspiration from biological learning, the speakers highlight that animals and humans learn primarily through experience rather than supervised learning with explicit targets. They emphasize the importance of forming abstractions and planning with learned models, capabilities currently missing in AI. The discussion touches on the challenge of enabling machines to perform paradigm shifts—creative leaps in understanding—through their own experience, a skill that distinguishes human intelligence. Sutton and Javed envision AI systems that can learn from the full spectrum of experience, from low-level sensory-motor skills to high-level abstract reasoning, all within a unified framework that maintains coherence and self-consistency over time.

Finally, the founders introduce their company, Oak Lab, which aims to realize this vision by developing algorithms for genuine continual deep learning and abstraction formation. They acknowledge the significant algorithmic challenges ahead and the need for a paradigm shift away from current large-scale, energy-intensive models toward more efficient, adaptive systems. Their goal is to build AI minds capable of lifelong learning and reasoning across diverse domains, not as a single monolithic entity but as a design instantiated in many agents with unique experiences. They seek a small, highly aligned team to pursue this ambitious agenda, confident that with the right algorithms and computational advances, truly intelligent, continually learning machines are within reach in the near future.