The Apple M5 Ultra chip offers substantial performance improvements over the M3 Ultra, including up to 4.3 times faster AI processing, doubled storage speeds, and enhanced memory bandwidth, making it especially powerful for AI developers and professionals handling large models and codebases. While delivering impressive speed gains, the M5 Ultra also consumes more power and generates more heat, positioning it as a high-end solution best suited for demanding workloads rather than casual use.
The M5 Ultra chip from Apple has arrived, boasting significant performance improvements over its predecessor, the M3 Ultra. Apple claims up to 4.3 times faster local AI performance, double the storage speed, and 1.3 times faster CPU speeds. The reviewer tested a high-end configuration with an 80-core GPU, 256 GB of memory, and 8 TB of storage, which costs around $14,299. Both the M3 Ultra and M5 Ultra were tested using identical software builds to ensure a fair comparison, revealing that the M5 Ultra delivers notably faster performance in various developer tasks, including compilation and Python-based algorithms.
Storage performance on the M5 Ultra is impressive, with sequential read and write speeds far exceeding those of the M3 Ultra, making it particularly beneficial for AI workloads that involve loading large models. The reviewer measured memory bandwidth using benchmarks and found that the M5 Ultra’s GPU memory bandwidth is about 1.4 times higher than the M3 Ultra, aligning well with Apple’s specifications. This increased bandwidth directly translates to faster token generation speeds in local large language models (LLMs), with the M5 Ultra outperforming the M3 Ultra by roughly 1.5 times in this area.
A key architectural improvement in the M5 Ultra is the integration of neural accelerators within each GPU core, enabling more efficient AI computations. This results in a dramatic increase in prompt processing speeds—up to four times faster than the M3 Ultra—especially noticeable with longer prompts or larger codebases. However, for shorter prompts typical of casual chatting, the speed gains are more modest. The M5 Ultra also consumes more power and generates more heat under heavy workloads, reflecting its higher performance capabilities but also requiring better cooling solutions.
The reviewer highlights that while the M5 Ultra is a powerful machine, it may be overkill for developers who do not require such extreme performance. The improvements are most beneficial for professional users working with large AI models or extensive codebases. The reviewer plans to conduct further tests, including comparisons between different AI software stacks and exploring the capabilities of even larger memory configurations, such as the upcoming 512 GB models.
In conclusion, the M5 Ultra delivers substantial upgrades in AI performance, storage speed, and memory bandwidth, validating Apple’s claims. Token generation is about 1.5 times faster, and prompt processing can be up to four times faster, making it a compelling choice for AI developers and professionals. However, these gains come with increased power consumption and heat output. The reviewer encourages viewers to subscribe for more in-depth testing and comparisons, emphasizing that the M5 Ultra represents a significant step forward in Apple’s chip technology for demanding AI workloads.
Useful Links
- Llama CPP - Cross-platform LLM library — Directly related to the software used for benchmarking AI performance on the M5 Ultra and M3 Ultra chips.
- STREAM Benchmark - Sustained Memory Bandwidth Measurement — Used to substantiate claims about memory bandwidth improvements on the M5 Ultra chip.
- Apple M5 Ultra Chip Official Product Page — Primary product discussed and benchmarked in the video.
- Deepseek V4 Flash Model - 284 Billion Parameter LLM — Used to demonstrate token generation performance improvements on the M5 Ultra.
- GPT OSS 120B Model — Benchmark model used to illustrate prompt processing improvements on the M5 Ultra.