Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
jialinyyzz 
posted an update 3 days ago
Post
4831
For a single-purpose 12B rewriter, how much loss per bit? Chinese breaks first.

We used llama.cpp's --kl-divergence for jialinyyzz/humanizer against bf16 weights, on our held out eval (drafts+rewrites) for English and Chinese. Standard llama-quantize (but our imatrix) without additional training. "differ" means top-choice token is not the same as what bf16 selects.

- Q6_K, 10.0 GB: KL .0030 EN / .0027 ZH. About 2 in 100 tokens differ in both.
- Q4_K_M, 7.6 GB: .0198 / .0204. About 6 in 100.
- IQ3_XXS, 4.7 GB: .138 / .175. 14 vs. 17 in 100.
- IQ2_XS, 3.8 GB: .487 / .762. 26 vs. 35 in 100.

Both languages track when you go as low as 4 bits. Below that Chinese drops off faster (Chinese KL is 1.3x English at 3 bits and 1.6x at 2 bits). Not clear why. One clue: when we added more Chinese to an imatrix (1/3 of the 800k tokens), it reduced Chinese KL by 4.7% at 2 bits. English didn't change.

(The ruler: These were running at about 8k tokens for each language. A 30690 token run agrees with English, but suggests Chinese was ~10% undercounted for 4 and 3 bits. So a bit more difference)

WIP/not released. We're distilling only the fp parts of the GGUF (block scales and norms) vs. bf16. Freezing integer codes. At 2 bits we're seeing approx. 1/2 KL (so far .487 -> .263 EN, .762 -> .319 ZH; different tokens 26 -> 20, 35 -> 22 / 100). Barely changed at 4 bits so we stopped there. No new stuff to grab yet.

The Chinese gap looks like it lives in the block scales, not the integer codes. Your own two rows say so.

Before the scale distillation, ZH/EN KL at 2 bits is 1.56 (.762 over .487). After it, with the codes frozen, 1.21 (.319 over .263). Top-1 flips go from 35 vs 26 to 22 vs 20. Chinese recovered 2.4x, English 1.85x, and the only thing that moved was scales and norms.

That makes the imatrix result the odd one out. A third of the calibration tokens in Chinese bought 4.7% at 2 bits. The imatrix only reweights the per-block fit in weight space; the scale still minimizes weighted weight error, not output KL. Distillation moves the same scales against KL directly, and Chinese gains more from that than English does.

One check before settling on "not clear why": bin tokens by the bf16 top-1 probability and plot KL per bin for each language. If the curves overlap, Chinese is not being quantized worse. It has more near-ties, and the same weight noise flips more of them. The Q6_K row already hints at this: 2 in 100 top-1 flips at KL .003 only happens where the top two choices are close.

What was the EN/ZH mix in the distillation set?

·

Nice read, thanks for the detailed analysis. For the distilled set it was about 72/28 for EN/ZH: 25.9k EN/10.2k ZH docs, 12.6M/4.8M tokens, so the minority ZH picked up more.

After this I have a result that partially confirms and partially complicates this: Training of the codes themselves (layer wise QAT, leaving scales to next stage) also favoured chinese more on Q2_K: ZH .483 --> .259 vs EN .354 --> .238, so ZH/EN changed from 1.36 to 1.09. So there is part of the gap also in codes, not just scales. Adding mixed precision + codes training + scale distillation now, for 3.9 GB build I'm at .106 / .106, so that's when the gap completely disappears.

I haven't checked the binning yet. It's a great idea and cheap to do. Will post curves once I have them.

That settles it, and against my guess. Near-ties are 18.8% vs 20.4%, and 94% of the gap sits inside the bins. I had it wrong.

The p1 >= .99 bin is the one I would keep. A flip there is not a paraphrase. bf16 was sure, so the token was close to forced: a copied number, a name, the second character of a word. Raw 2-bit flips 6.2% of those in Chinese and 0.4% in English, about 15x.

The new fact column on the card is the other half of this, and I read it a little differently than footnote 3 does. 2-bit against Q8_0: English goes from 44 to 70 flagged of 420, called beyond noise. Chinese goes from 54 to 66 of 204, called within noise. In points that is +6.2 and +5.9. Same shift. Chinese has half the rewrites, so the same effect cannot clear the bar.

So on facts the 3.9 GB build costs both languages about 6 points, on top of a Chinese baseline that is already 26.5% flagged against 10.5%. The fix closed the KL gap and did not open a fact gap. It did not leave Chinese untouched either.

Do the paired counts say the same? How many Chinese rewrites were flagged only at 2-bit, and how many only at Q8_0?