Can A Computer With No GPU Run A Million-Token AI?

The Spark X 2.5 4B model can run on CPUs without GPUs and supports a massive one million token context window, but processing such large contexts is extremely slow and memory-intensive, making real-time use impractical on typical hardware. For effective performance without a GPU, it is recommended to limit the context window to tens of thousands of tokens and disable the model’s “thinking mode” to avoid long delays and potential errors.

The Spark X 2.5 4B model, developed by iFLYTEK’s SparkLLM team, is notable for its native support of a massive one million token context window, allowing it to process text equivalent to several thick books simultaneously. The model is available in a compressed 4-bit version of about 2.6 GB, making it accessible for on-device use. A token represents a word or part of a word, and the context window defines how much text the model can consider at once. While the model can technically run on ordinary computers without a graphics processing unit (GPU), fully utilizing the million-token window presents significant challenges.

Testing has confirmed that the model boots and responds to prompts on CPUs without GPUs, including an eight-core server with 16 GB of shared memory and an 18-core ARM machine. However, the speed is very slow, with writing speeds ranging from six to 47 tokens per second depending on the hardware. Reading speeds improve with more processor threads but writing performance can degrade if too many threads compete for resources. Users must also manually adjust the context window size, as the default setting attempts to use the full million tokens, which can overwhelm systems with limited memory.

Memory requirements for holding a million tokens are substantial but manageable on desktops with 64 GB of RAM. The model uses a KV cache to store token information, which consumes about 39 GB for a million tokens plus the model size itself. This efficiency is achieved by limiting full attention caching to nine of the 36 layers, with the rest only caching the latest 512 tokens. Systems with less RAM, such as 32 GB or 16 GB, can only handle smaller context windows or require compression, which may affect accuracy. The model’s design helps keep memory demands within reach for high-end consumer hardware without GPUs.

The main bottleneck is processing time. Reading a million tokens on a CPU can take around 13 hours on a fast 18-core ARM machine and up to four days on an eight-core Xeon server. This is due to the quadratic increase in computation as the model compares each new token against all previous tokens in the full-attention layers. Even smaller prompts of tens of thousands of tokens take minutes to process, making real-time interaction impractical at very large context sizes. GPUs significantly speed up this process, with a small graphics card able to read over 130,000 tokens in about four minutes, highlighting the advantage of hardware acceleration.

Finally, while the model performs well up to around 100,000 tokens, with benchmark scores comparable to similar models, there is no published evidence supporting reliable performance beyond that scale. Additionally, the model’s default “thinking mode,” which generates internal reasoning steps, can drastically increase response times and cause it to get stuck in loops, especially on CPUs. For practical use without a GPU, it is recommended to limit the context window to tens of thousands of tokens and disable thinking mode for faster, more reliable responses. In summary, running Spark X 2.5 4B on a CPU without a GPU is feasible but using the full million-token window is hindered by long wait times and uncertain output quality at extreme lengths.

Useful Links