1.58 bit LLMs Are BETTER, But Why No One Uses Them?

In 2024, Microsoft introduced 1.58-bit ternary quantization for large language models, significantly reducing computational costs and memory usage while maintaining performance comparable to 16-bit models, enabling efficient operation on consumer CPUs without GPUs. However, challenges such as expensive retraining, limited information capacity, hardware incompatibility, and industry resistance have hindered widespread adoption despite the method’s potential to democratize AI access.

In 2024, Microsoft researchers introduced a groundbreaking approach called Bitnet that drastically reduces the computational demands of large language models (LLMs) by replacing complex multiplications with simple additions and subtractions. This method, known as ternary quantization, represents model weights using only three states (+1, 0, -1) combined with a shared scale factor, effectively compressing weights to about 1.58 bits each. The result is models that perform comparably to full-precision 16-bit models on benchmarks while consuming up to 82% less power, using 75% less memory, and running efficiently on consumer CPUs without GPUs. Despite these impressive gains, major AI players like OpenAI, Anthropic, and Google have yet to adopt this technology widely.

The core challenge lies in training these ternary models. Unlike simpler 4-bit quantization, which approximates weights to 16 levels, ternary quantization drastically reduces precision to just three levels, causing significant information loss if applied directly to pre-trained models. To overcome this, researchers use quantization-aware training, where two versions of weights are maintained: a full-precision copy for backpropagation and a ternary version for forward passes. This method, called the straight-through estimator, allows the model to learn weights that naturally round to ternary values. However, training remains computationally expensive, requiring extensive retraining from scratch or large-scale fine-tuning, which has limited adoption.

Another limitation of ternary models is their capped information capacity due to the 1.58-bit constraint, which affects their ability to memorize facts. For example, Microsoft’s 2-billion parameter ternary model performs well on reasoning tasks but falls short on factual recall compared to larger 16-bit models. This trade-off highlights that while ternary models excel in efficiency and reasoning, they struggle with knowledge retrieval, akin to how hard drives store data precisely but retrieval can be cumbersome without exact queries. Efforts like Prism ML’s Bonsai 2 model demonstrate progress by fine-tuning ternary models from pre-trained full-precision teachers, achieving near-original performance with drastically reduced memory footprints.

Hardware compatibility also poses a significant hurdle. Modern GPUs are optimized for multiplication-heavy operations using tensor cores, which are inefficient for the addition/subtraction operations ternary models require. Consequently, ternary weights must be unpacked back into full precision before multiplication, negating some efficiency gains. Meanwhile, Nvidia’s new Reuben GPUs support 4-bit floating-point operations natively, offering a practical compromise that reduces memory usage without requiring radical changes to hardware or software. Given Nvidia’s dominance in AI hardware, their preferred formats heavily influence industry standards, making widespread adoption of ternary models less likely in the near term.

In summary, 1.58-bit ternary LLMs offer remarkable efficiency improvements and maintain much of the intelligence of full-precision models, making them ideal for users without access to expensive GPUs. However, the high cost of retraining, limited memory capacity, hardware incompatibilities, and industry inertia have slowed their adoption. While ternary models represent a promising direction for democratizing AI by enabling powerful models on consumer hardware, overcoming these challenges will be crucial before they become mainstream.

Useful Links