Back to research

PaperML hardwarePreprint

Natural-Sparsity Acceleration for Ternary LLMs

Ternary language models store each weight as −1, 0, or +1. This paper asks whether the zeros already in a trained model are worth a different format and chip design, without retraining: a bitmap of which weights are nonzero, plus a packed list of their signs, instead of the usual five-trit packing. In simple terms, five-trit packing stores every weight the same way; the bitmap skips the zeros and keeps only the signs of the rest. Matched matrix–vector engines for both formats were compared in the same RTL, timing-aware synthesis, and gate-level energy flow in Nangate45, including the logic that reads weights from memory. The bitmap engines reach a 9–14% shorter clock, but they are larger and use more energy in the compute path, because five-trit decoding already drops zero weights before the shared adder tree. The bitmap stores fewer bits only above 40% zeros, so any real saving is from reading less memory, not from doing less math. Across 13 public checkpoints, nine quantization-aware or fine-tuned models (29–42% zeros) save at most 1.3% in the modeled projection path. The sparser CAT-Q Qwen3 models save 3.7–6.8% when ternary weights stream from LPDDR4, falling to about 1.9–2.1% once scale and output-head reads are included. The natural zeros help only when a model is unusually sparse and fetching weights is expensive. They are not a general reason to replace five-trit packing.