BITCOS: Intel's Ternary LLM Weight Format Goes Below 1.58 Bits
If you ship or run ternary models, do not treat 1.58 bits as the floor. Measure each tensor first. Above 37.5% zeros, a presence bitmap plus a sign vector is smaller than the five-trits-per-byte packing that kernels use today. On Microsoft's BitNet b1.58 2B4T, the tensors I measured have 38.5% to 50.8% zeros. A smaller file is faster only when you have enough cores to unpack it. On a model with tied embeddings, the higher-precision output head is larger than all the ternary layers together.
The layout comes from Breaking the 1.58-bit Barrier for Ternary LLMs, posted September 14 by three Intel researchers. It stores a real ternary checkpoint at 1.485 bits per weight, with no retraining and no lost values. This works because 29.66% to 51.48% of all weights are zero across 29 published ternary models. The log2(3) figure holds only when the three values are equally common. Two days later the paper had 245 points on Hacker News. This essay covers how the encoding works, code to measure and pack your own checkpoint, and where the speedup stops.
What BITCOS is. BITCOS, short for bitmap and compacted signs, is a weight-packing format for ternary large language models. Intel Labs researchers Evangelos Georganas, Alexander Heinecke, and Pradeep Dubey introduced it. It is a file and memory layout for model weights, with CPU and GPU unpacking kernels to match. It has nothing to do with the cryptocurrency or DeFi token that shares the name.
Where 1.58 comes from
Start with the ladder most people already know. A weight in FP16 costs 16 bits. INT8 costs 8. The 4-bit formats that run most local models cost 4 bits, plus a small share for the per-group scales.
Each step down halves the bytes that cross the memory bus for every generated token. I wrote about this in Tokens per Second Is a Memory Bandwidth Number. Single-stream decode speed follows the size of the weights more closely than FLOPs.
Ternary models are at the far end of that ladder, and they are a different kind of object. You cannot take a trained FP16 model and round it to three values without destroying it. You must train the model, or fine-tune it heavily, with ternary weights in the loop. Microsoft's BitNet b1.58 paper from February 2024 made the recipe popular. Scale a weight matrix by its mean absolute value, round, and clip to the set {-1, 0, +1}.
The "1.58" in the name is log2(3). Suppose each weight is one of three symbols and all three are equally likely. Then information theory says you need log2(3) bits per weight on average, and no encoding can do better.
The new paper questions the "equally likely" part. Shannon's bound is the entropy of the actual distribution of symbols. That entropy equals log2(3) only when the distribution is flat. When the mix changes, the floor moves.
What the shipping packers actually store
Deployed kernels do not store 1.585 bits per weight either. The paper describes the two common layouts.
The simplest is 2 bits per weight. Two bits hold four states, three of which are used, and decoding is a shift and a mask. It wastes a quarter of the space, but it is the fastest layout to unpack. The paper uses 2-bit kernels as its speed baseline for that reason.
The tighter one packs five ternary digits into a byte. Five trits have 35 = 243 combinations, which fits under 256, so the rate is 8/5 = 1.6 bits per weight. That is close to the bound. But real kernels quantize in blocks of a power of two, typically 128 weights, and 5 does not divide 128. A block needs ⌈128/5⌉ = 26 bytes, so the paper puts the stored rate at 26 × 8 / 128 = 1.625 bits per weight. BITCOS has to beat that rate.
A block of 128 ternary weights, the size real kernels quantize in. Pick how many are zero, or load a real model's measured zero density, then switch encodings. Each cell is one weight: filled blue is +1, ink is -1, hollow is 0. The strips underneath are the bits the encoding writes for this exact block.
Block layout and the 1.625 bits per weight for five-trit packing at a 128-weight block are from the paper. A fixed seed generates the tile with the chosen zero count and balanced signs. No encoding includes the per-block scales.
The zeros nobody was counting
The authors measured the zero density of 29 published ternary checkpoints across seven families. The lowest is Bonsai 27B at 29.66%. The highest is CAT-Q Qwen3-1.7B at 51.48%. BitNet b1.58 2B4T, the model most people have run, sits at 42.19%. The other models fall in between:
- the other Bonsai models, 37.71% to 39.89%
- the TriLM family, 38.70% to 40.97%
- the BitCPM-CANN models, 37.67% to 39.30%
- Maple 20B-A1B, 40.67%
- the ParetoQ models, 41.07% to 47.10%
- the other CAT-Q Qwen3 conversions, 32.88% to 47.11%
A flat distribution puts each symbol at 33.3%. Every model in the set except Bonsai 27B has more zeros than that, and several have many more.
The quantizer explains why. The BitNet recipe divides each weight by the matrix's mean absolute value and rounds. Any weight within half a mean absolute value of zero becomes zero.
If the pre-quantization weights are roughly Gaussian, that window catches about 31% of them. That is my arithmetic from the recipe, not a number from the paper. It is close to the least sparse model in the survey. In most families, training with the quantizer in the loop pushes more weights toward zero, and that gives the 40% and 50% figures.
You can measure your own model instead of trusting the survey. The zero density comes from the quantizer, so you can compute it from the higher-precision master weights. You do not have to run the model. Per its technical report, BitNet b1.58 2B4T quantizes with an absmean rule across each weight tensor. Microsoft publishes the bf16 master weights as microsoft/bitnet-b1.58-2B-4T-bf16.
The script below uses only the Python standard library. It reads the safetensors header with HTTP range requests and pulls one tensor at a time. Then it applies the absmean rule and counts zeros. A weight rounds to zero when its magnitude is less than half the tensor's mean absolute value.
import json, struct, sys, urllib.request
from array import array
URL = "https://huggingface.co/microsoft/bitnet-b1.58-2B-4T-bf16/resolve/main/model.safetensors"
def fetch(start, end):
buf = bytearray()
while start + len(buf) <= end:
req = urllib.request.Request(URL, headers={"Range": f"bytes={start + len(buf)}-{end}"})
try:
with urllib.request.urlopen(req) as r:
while part := r.read(1 << 16):
buf += part
except Exception:
pass # resume from what arrived
return bytes(buf[: end - start + 1])
def header():
n = struct.unpack("<Q", fetch(0, 7))[0]
return 8 + n, json.loads(fetch(8, 8 + n - 1))
def bf16_to_f32(raw):
out = bytearray(len(raw) * 2)
out[2::4], out[3::4] = raw[0::2], raw[1::2]
return array("f", bytes(out))
base, h = header()
for name in sys.argv[1:]:
a, b = h[name]["data_offsets"]
w = bf16_to_f32(fetch(base + a, base + b - 1))
cut = 0.5 * sum(map(abs, w)) / len(w)
z = sum(1 for x in w if -cut < x < cut) / len(w)
print(f"{name} z={z:.4f} bitcos={2 - z:.4f} bits/w")
Run as python3 zero_density.py model.layers.0.self_attn.q_proj.weight .... The resume loop exists because some proxies cut long range reads a few kilobytes short. I ran the script on September 25 against ten of the model's 210 ternary tensors:
| Tensor | Shape | Zeros | BITCOS bits/w | vs 1.625 |
|---|---|---|---|---|
| layers.0.self_attn.q_proj | 2560 x 2560 | 50.83% | 1.4917 | smaller |
| layers.0.self_attn.o_proj | 2560 x 2560 | 47.81% | 1.5219 | smaller |
| layers.0.self_attn.k_proj | 640 x 2560 | 45.84% | 1.5416 | smaller |
| layers.15.self_attn.q_proj | 2560 x 2560 | 44.59% | 1.5541 | smaller |
| layers.15.mlp.gate_proj | 6912 x 2560 | 44.14% | 1.5586 | smaller |
| layers.0.self_attn.v_proj | 640 x 2560 | 41.72% | 1.5828 | smaller |
| layers.29.mlp.down_proj | 2560 x 6912 | 40.84% | 1.5916 | smaller |
| layers.0.mlp.gate_proj | 6912 x 2560 | 38.93% | 1.6107 | barely |
| layers.0.mlp.down_proj | 2560 x 6912 | 38.79% | 1.6121 | barely |
| layers.0.mlp.up_proj | 6912 x 2560 | 38.54% | 1.6146 | barely |
The spread inside one model, 38.5% to 50.8%, is nearly as wide as the spread across the paper's 29 models. So the paper's 42.19% for this model is an average over tensors that behave quite differently. The first layer's MLP is just above the break-even line. Its attention projections are near the paper's best case. A converter that picks one layout for the whole file wastes bytes on both kinds of tensor.
A bitmap and a sign
BITCOS splits every weight into two questions and stores the answers in two separate arrays. The first is a presence bitmap: one bit per weight, set when the weight is non-zero. The second is a sign vector: one bit per non-zero weight only, in the same order, saying whether it is +1 or -1. Zeros cost their one bitmap bit and nothing else.
The cost follows directly. With a zero density of z, every weight pays 1 bit for the bitmap. A fraction (1 − z) of the weights pay 1 more bit for the sign. So the rate is 1 + (1 − z) = 2 − z bits per weight. The paper writes it the same way. At CAT-Q Qwen3-1.7B's 51.48%, that is 1.4852, the 1.485 in the headline.
The formula also shows when BITCOS loses. It beats 1.625-bit five-trit packing when 2 − z < 1.625, which is any zero density above 37.5%. Three models in the survey are below that line: Bonsai 27B at 29.66%, CAT-Q Qwen3-30B-A3B at 32.88%, and CAT-Q Qwen3-235B-A22B at 34.07%.
The paper reports that BITCOS wins in 26 of 29 models. The three losses are the three that the formula predicts. They are the two largest mixture-of-experts conversions and the largest Bonsai. Bonsai 4B, at 37.71%, is above the line by two hundredths of a bit.
The layout fits in a few lines of code, and writing it is the fastest way to check that the formula is exact. This packer and unpacker work on a flat list of ternary codes. The round-trip assertion is the test.
def pack_bitcos(q):
bitmap = bytearray((len(q) + 7) // 8)
signs, acc, nbits = bytearray(), 0, 0
for i, v in enumerate(q):
if v:
bitmap[i >> 3] |= 1 << (i & 7) # presence bit
acc |= (v < 0) << nbits # sign bit, non-zeros only
nbits += 1
if nbits == 8:
signs.append(acc); acc = nbits = 0
if nbits:
signs.append(acc)
return bytes(bitmap), bytes(signs)
def unpack_bitcos(bitmap, signs, n):
out, k = [], 0
for i in range(n):
if bitmap[i >> 3] >> (i & 7) & 1:
out.append(-1 if signs[k >> 3] >> (k & 7) & 1 else 1); k += 1
else:
out.append(0)
return out
bitmap, signs = pack_bitcos(q)
assert unpack_bitcos(bitmap, signs, len(q)) == q
Ours, built from the layout that the paper describes. It writes one presence bit per weight and one sign bit per non-zero weight, in order. A production kernel also aligns both arrays to its block size.
On two real tensors from the table above, the packed bytes land exactly on 2 − z:
layers.0.self_attn.q_proj 6,553,600 weights, 50.83% zeros
2-bit 1.638 MB 2.0000 bits/w
five-trit (26 B per 128) 1.331 MB 1.6250 bits/w
BITCOS bitmap + signs 1.222 MB 1.4917 bits/w
layers.0.mlp.up_proj 17,694,720 weights, 38.54% zeros
2-bit 4.424 MB 2.0000 bits/w
five-trit (26 B per 128) 3.594 MB 1.6250 bits/w
BITCOS bitmap + signs 3.571 MB 1.6146 bits/w
So the choice per tensor is one line of code. Use BITCOS where it is smaller and five-trit packing where it is not. Record the choice in the file so that the loader calls the correct decoder:
def choose_layout(z):
return "bitcos" if 2 - z < 1.625 else "five-trit" # break-even at z = 0.375
Our size rule, derived from the paper's two rates. It says nothing about speed. The next section covers speed.
How close does 2 − z get to the real floor? Assume signs are balanced between +1 and −1. Then the entropy of the ternary distribution is h(z) + (1 − z). Here h(z) is the binary entropy of the question "is this weight zero."
BITCOS spends exactly 1 bit on that question. The value h(z) is at most 1, when z = 0.5. So the gap between BITCOS and the Shannon floor is 1 − h(z). At z = 0.5148, that gap is 0.0006 bits per weight. For the sparsest model in the survey, the layout is within a rounding error of optimal.
One Hacker News commenter said they wanted to write a follow-up paper with arithmetic coding. The aim was "to squeeze out a few more centi-bits."
The formula gives the size of that gain. At a zero density of 40%, an ideal entropy coder saves 0.029 bits per weight over BITCOS. At 30%, it saves 0.12. So "a few centi-bits" is correct. Each of those bits also costs decode work, and decode work is the subject of the next section.
The ruler runs from 1.4 to 2.1 bits per weight. Drag the zero density, or tap a model from the paper's survey, and watch where the Shannon floor, BITCOS, five-trit packing, and the 2-bit layout land. The sign skew slider unbalances +1 against −1, which lowers the floor but not BITCOS, since BITCOS always spends one full bit per sign.
Zero densities are the paper's measurements for individually named models. Shannon floor = h(z) + (1 − z)·h(s) with s the +1 share. The speedup line is the ideal bytes ratio 2 / (2 − z) against a 2-bit kernel. Measured speedups are lower.
Decoding is the other half of the bill
A smaller file helps only if the kernel unpacks it faster than the memory bus delivers it. Every decode token streams the whole weight matrix through the compute units once. So the maximum gain from BITCOS is the byte ratio against the 2-bit baseline, 2 / (2 − z). At 51.48% zeros, that ceiling is 1.35x. At 29.66%, it is 1.17x. All unpacking work comes out of that margin.
The paper's decoders keep that work small. On CPUs with AVX-512, the kernel loads 32 bits of the presence bitmap into a mask register. Then pdep, the parallel bit deposit instruction, puts the next compacted sign bits into the positions that the mask marks as present. Every unmarked position comes out zero. Masked moves then select +scale or −scale per lane.
AVX2 has no mask registers. On AVX2, the kernel expands presence and sign into byte masks and computes each weight as p − 2n, presence minus twice the negative flag. On Intel's Xe2 GPUs, it uses a 256-entry lookup table in shared local memory. A 4-bit presence nibble and a 4-bit window of sign bits index the table, and each lookup returns four fp16 ternary codes.
The measured results stay under that ceiling. Across the survey's range of zero densities, matrix-vector multiplication was faster than the 2-bit reference kernel by these amounts:
- 64-core Emerald Rapids Xeon: 1.14 to 1.28x
- 24-core Arrow Lake desktop part: 1.13 to 1.27x
- Arc 140V integrated GPU: 1.04 to 1.14x
- Arc Pro B70 discrete card: 1.01 to 1.12x
End to end, decode throughput improved by 1.10 to 1.18x on Emerald Rapids and by 1.02 to 1.27x on the B70. The 1.28x top result on the Xeon is close to the 1.35x byte ceiling. So on that chip, the AVX-512 decoder costs almost nothing.
| Platform | Bandwidth | GEMV vs 2-bit | End-to-end decode |
|---|---|---|---|
| Emerald Rapids, 64 cores | ~245 GB/s | 1.14 to 1.28x | 1.10 to 1.18x |
| Arrow Lake, 24 cores | ~98 GB/s | 1.13 to 1.27x | 1.02 to 1.15x |
| Arc 140V (integrated) | ~108 GB/s | 1.04 to 1.14x | 1.09 to 1.22x |
| Arc Pro B70 (discrete) | ~500 GB/s | 1.01 to 1.12x | 1.02 to 1.27x |
| Lunar Lake, 8 cores | ~108 GB/s | slower at every density | not a win |
The Lunar Lake row is the most useful result in the paper. That chip has 108 GB/s of bandwidth and only eight cores. The authors measured 3.2 to 3.7 bytes per cycle of bandwidth available to each core. Through the BITCOS unpack sequence, each core consumed only 0.87 to 1.38 bytes per cycle. So the kernel sustained 28.9 GB/s at a zero density of 40%, and it lost to the simple 2-bit kernel at every density.
When bandwidth is high and cores are few, the cheapest decode wins, even if it moves more bytes. The paper reports the same effect at the other extreme. Near 95% zeros, the unpack instructions become the bottleneck on every platform.
So the rule for local inference is the one from the memory-bandwidth essay, with a second term. Tokens per second is bytes moved divided by bandwidth, as long as the decoder keeps up. BITCOS shrinks the bytes. The smaller size becomes speed only when enough cores are available to do the unpacking.
Two checks tell you which side of that line your machine is on. First, check whether the fast decoder can run at all. The paper's best CPU path needs AVX-512 mask registers and pdep, which Intel ships as part of BMI2. Without them, it falls back to an AVX2 sequence.
# Linux: which of the paper's decode paths this CPU can take
lscpu | grep -o 'avx512f\|avx2\|bmi2' | sort -u
Second, measure your current ternary runtime, so that any new kernel has a number to beat. Microsoft's bitnet.cpp documents this path for the same model:
huggingface-cli download microsoft/BitNet-b1.58-2B-4T-gguf --local-dir models/BitNet-b1.58-2B-4T
python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s
python utils/e2e_benchmark.py -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf -n 200 -p 256 -t 4
From the bitnet.cpp README, not run here. Repeat the benchmark with -t set to 2, 4, and your core count. If tokens per second increases with each thread count, decode is compute-bound on that machine. Then a heavier unpack sequence makes it slower.
Worked example: where the bytes go in BitNet b1.58 2B4T
Take the model from the measurements above and count the bytes that one decode step reads. The safetensors header gives the inputs. There are 30 layers of seven ternary projections, 2,084,044,800 ternary weights in total. There is also a token embedding of 128,256 x 2,560 = 328,335,360 parameters. The config sets tie_word_embeddings to true, so that embedding is also the output head. On every generated token, the head multiplies against the whole matrix.
| Per-token read | 2-bit layers | Five-trit layers | BITCOS layers (z = 0.4219) |
|---|---|---|---|
| Ternary layers | 521.0 MB | 423.3 MB | 411.1 MB |
| + head in f16 | 1,177.7 MB | 1,080.0 MB | 1,067.8 MB |
| + head in q8_0 | 869.9 MB | 772.2 MB | 760.0 MB |
| Ceiling at 100 GB/s, f16 head | 84.9 tok/s | 92.6 tok/s | 93.7 tok/s |
| Ceiling at 100 GB/s, q8_0 head | 115.0 tok/s | 129.5 tok/s | 131.6 tok/s |
Our arithmetic uses weights x bits / 8 and the paper's 42.19% model-level zero density. It counts f16 at 2 bytes and q8_0 at 8.5 bits per parameter. The ceiling is bandwidth over bytes, which assumes that the decoder keeps up. Norms and the KV cache are not included.
Against five-trit packing, BITCOS saves 12.2 MB per token on this model. bitnet.cpp's own --quant-embd option keeps the tied head in f16. That costs 307.8 MB more per token than q8_0. So fix the head first. llama.cpp's quantizer has a direct option for it, and bitnet.cpp's setup script already calls the same flag with f16:
./build/bin/llama-quantize --token-embedding-type q8_0 \
models/BitNet-b1.58-2B-4T/ggml-model-f32.gguf models/BitNet-b1.58-2B-4T/ggml-model-i2_s-q8head.gguf I2_S 1 1
The flag is documented in the llama.cpp quantize README. The argument order comes from bitnet.cpp's setup_env.py, which passes f16 in that position. Not run here. Check output quality on your own prompts before you keep it.
Next, check the core count. On the paper's eight-core Lunar Lake laptop, the BITCOS kernel sustained 28.9 GB/s at 40% zeros, against 108 GB/s of memory bandwidth. At that decode rate, the ternary layers alone limit decode to about 69 tokens per second. A 2-bit kernel that keeps pace with the memory bus has a ceiling near 207. That gap explains why the simple layout won on that chip at every density. The saved bytes become speed only on machines with enough cores for unpacking.
When it goes wrong
| Symptom | Cause | Fix |
|---|---|---|
| Packed file is larger than the five-trit build | Tensors under 37.5% zeros, like Bonsai 27B or the large CAT-Q mixture-of-experts conversions | Choose the layout per tensor with choose_layout(z) and keep five-trit where it wins |
| Smaller file, slower tokens | Too few cores to unpack at memory speed, as on the eight-core Lunar Lake part | Benchmark across thread counts and keep the 2-bit kernel on that machine |
| Barely any gain end to end | A tied or higher-precision head dominates the bytes per token | Quantize the head first. It is often worth more than the whole ternary re-pack |
| Your zero share disagrees with a published number | Different scaling rule, per tensor versus per row or per group, or a converted checkpoint | Measure with the rule the model was trained with, on the weights you will ship |
| Decode stalls on very sparse models | Near 95% zeros the unpack instructions, not memory, become the bottleneck on every platform the paper tested | Treat extreme sparsity as a kernel problem. Do not expect the full byte ratio as speed |
What changes, and what stays the same
BITCOS is lossless. It does not change a single weight. It re-encodes the ternary values that a model already has. So it needs no retraining and no special sparsity hardware. The paper contrasts it with approaches like Sparse-BitNet, which train in structured sparsity.
You can convert any ternary checkpoint. A converter can compute 2 − z for each checkpoint and use five-trit packing when z is under 37.5%. That per-model choice is my suggestion, not something that the paper ships.
The memory savings are real but modest. In the best case, the gain over five-trit packing is 1.625 − 1.485 = 0.14 bits per weight, about 8.6% smaller. For an 8-billion-parameter model, bits per weight times 8 billion divided by 8 bits per byte equals gigabytes. So 1.625 GB of ternary weights becomes 1.485 GB. That number does not include the embeddings, norms, and scales that stay in higher precision.
Whether ternary models matter at all is a separate question, and the Hacker News thread disagreed about it. One commenter said that vector quantization and trellis-based methods beat ternary for post-training quantization at these bit rates. Another pointed at hardware. A format with only a bitmap and a sign is easy for a small accelerator to read directly.
Ternary checkpoints still have to be trained as ternary. In dense form, the families in the survey stay well below frontier scale.
For a real checkpoint, the mechanism gives a short order of work:
- Measure the zero share of every ternary tensor with the model's own quantization rule.
- Compare 2 − z against 1.625 for each tensor. Use the smaller layout.
- Before you re-pack anything, find what else a decode step reads, especially the output head.
- Shrink the largest item first.
- Make sure that the target machine has the cores and instructions to unpack faster than its memory bus.
- Benchmark against your current kernel at several thread counts.
- Keep the fallback layout wherever it wins.
The paper settles one narrower point. The information floor for a ternary model comes from that model's weights, not from the number three. Measure the distribution before you choose the container. For the sparsest published checkpoints, that measurement leads to two bit arrays and the formula 2 − z. The result is within six ten-thousandths of a bit of the Shannon floor.
Keep reading