The Inference Inflection from First Principles — swyx & Rob Wachen, Etched

Rob Wachen, co-founder of Edge, discusses their innovative AI inference chip and cluster technology that integrates hardware and system design to overcome thermal and latency challenges, enabling efficient operation of large-scale AI models. Emphasizing a unique engineering culture and advanced low-latency interconnects, Edge aims to revolutionize AI inference infrastructure with scalable, high-performance solutions and invites talented engineers to join their mission.

The discussion features Rob Wachen, co-founder and president of Edge, who recently emerged from stealth mode with a groundbreaking AI inference chip and cluster technology. Rob explains that their journey began before the widespread adoption of ChatGPT, recognizing early on the immense market potential for AI inference hardware. Unlike traditional approaches focusing solely on chips, Edge vertically integrates the entire inference cluster, including chips, racks, cooling, power delivery, and interconnects, to meet the demanding requirements of serving trillion-parameter models with thousands of concurrent users and strict latency SLAs.

A key innovation highlighted is their approach to low voltage inference (LVI), which addresses the thermal throttling limitations common in existing AI accelerators. By redesigning power delivery and chip architecture, Edge enables chips to run at significantly lower voltages, reducing power consumption and allowing more cores to operate at full speed without overheating. This holistic co-design philosophy extends across materials, packaging, circuit design, and thermal management, demonstrating a deep integration of hardware and system engineering to maximize performance and efficiency.

Rob emphasizes the importance of building a unique engineering culture that fosters end-to-end ownership and cross-disciplinary collaboration. Unlike traditional semiconductor companies where teams work in silos, Edge’s engineers are responsible for architecture, design, testing, and production, enabling rapid iteration and problem-solving. Recruiting is focused on individuals who deeply identify with their work and are driven to solve hard problems, creating a team that is highly motivated and aligned with the company’s ambitious goals.

The company also tackles the challenge of cluster-scale memory and interconnects, aiming to build systems with tens of thousands of chips sharing memory with ultra-low latency. To achieve this, they are reinventing networking stacks and leveraging expertise from high-frequency trading to design custom, high-speed, low-latency interconnects that bypass traditional switches. This approach is critical for efficiently running large mixture-of-expert (MOE) models and represents a significant departure from existing GPU-centric hardware ecosystems.

Looking ahead, Rob shares that while their first-generation product is a major milestone, they view it as just the beginning, with subsequent generations already in development to push closer to theoretical performance limits. The company maintains a pragmatic yet visionary stance, focusing on shipping real hardware and iterating quickly. They invite talented engineers from diverse disciplines to join their mission to revolutionize AI inference infrastructure, underscoring the transformative potential of abundant, fast, and cheap AI inference in reshaping knowledge work and everyday applications.