The episode of The AI Hardware Show highlights the latest custom CPUs and AI accelerators developed by major cloud providers like Google, Amazon, Meta, and Microsoft, showcasing innovations such as Google’s Axian CPU, Amazon’s Inferentia 2, Meta’s MTIA V2, and Microsoft’s Cobalt 100 and Maya 100 chips, all designed to optimize AI inference and cloud workloads at hyperscale. It also touches on other significant players like Amazon’s Graviton 4 and China’s Baidu Kunlun 2 and Alibaba Honey Badger, emphasizing the strategic role of these specialized chips in advancing AI infrastructure and the need for greater transparency in Chinese semiconductor developments.
The episode of The AI Hardware Show explores the latest developments in hyperscale silicon, focusing on custom CPUs and AI accelerators designed by major cloud providers. Google’s Axian CPU, based on ARM’s Neoverse V2 architecture, is their first in-house general-purpose server CPU, delivering up to 72 single-threaded cores per socket. It features a unique offload layer called Titanium that handles networking, storage, and security, freeing CPU cores for user workloads. Axian supports AI inference workloads efficiently through ARM scalable vector extensions and brain float quantization, fitting seamlessly into Google’s broader silicon strategy alongside specialized tensor and TPU chips.
Amazon’s AWS Inferentia 2 is a dedicated AI inference chip designed for high throughput and low latency in production-scale deep learning models. Developed by Amazon’s internal Anaperna Labs, Inferentia 2 improves on its predecessor with increased memory, bandwidth, and software flexibility. It supports multiple precision formats and scales across chips to handle massive models. Integrated tightly with AWS’s AI infrastructure and SDKs, Inferentia 2 enables cost-effective deployment of AI workloads like natural language processing and recommendation systems, emphasizing efficiency and predictable performance at hyperscale.
Meta’s MTIA is a purpose-built inference accelerator tailored for ranking and recommendation models powering Facebook and its apps. The second-generation MTIA V2 chip, built on TSMC’s 5nm process, significantly boosts performance and efficiency, integrating custom RISC-V cores for fine-grained control and handling sparse data. It is tightly coupled with Meta’s PyTorch framework and Triton compiler, allowing seamless software integration without rewriting kernels. Deployed globally in Meta’s data centers, MTIA excels at latency-sensitive workloads with unpredictable batch sizes, though it remains an internal-use chip focused on operational efficiency and rapid iteration.
Microsoft has developed two key custom chips: the Cobalt 100 ARM-based CPU for general-purpose cloud workloads and the Maya 100 AI accelerator for large-scale AI tasks. Cobalt 100 offers up to 96 cores with improved performance and cost efficiency, supporting AI inference alongside traditional cloud applications. Maya 100, fabricated on TSMC’s 5nm process, features a large tensor core array, high-bandwidth memory, and a unique Ethernet-based chip-to-chip interconnect. It powers Azure OpenAI services and represents Microsoft’s vertically integrated approach to AI infrastructure, emphasizing open standards and scalable networking to challenge proprietary interconnects.
Finally, the episode covers Amazon’s Graviton 4 CPU, Baidu’s Kunlun 2 AI accelerator, and briefly mentions Chinese hyperscale chips like Alibaba’s “Honey Badger.” Graviton 4, with 96 ARM Neoverse V2 cores and chiplet design, excels in general cloud compute and AI-adjacent workloads. Kunlun 2, built on TSMC’s 7nm process, delivers high AI throughput and versatility across cloud and edge applications, supporting China’s push for semiconductor self-sufficiency. The discussion highlights the strategic importance of these custom chips in enabling hyperscale AI deployments, with a call for more transparency on Chinese designs.