Ok, ran the binning test. Basically, no, it's not a near-tie thing. Chinese really was more poorly quantized (in 2 bit).
In BF16, it's almost the same difficulty level for both languages. Output tokens have 18.8% (EN) 20.4% (ZH) near-ties (p1-p2 < .2). So it's hard to use the mix as an explanation.
In raw IQ2_XS, chinese kl is 1.3-6.4x english for all p1 bins. Most of the difference (94%) is within the bin. The differences are largest in places where bf16 is most confident (p1 >= .99) - chinese top1 flips 6.2% (vs 0.4% for english). This is the reverse of what you would expect for the near-tie thing.
In Q4_K_M, Q6_K, and new 3.9 GB 2-bit, the curves are roughly aligned in low/mid bins. The remaining difference is about 8% and primarily in p1 >= .9 bins (very low absolute kl). 2-bit reduces output token kl from .315/.573 (EN/ZH) to .056/.060.
So yes, raw low-bit quant hurts chinese more. But it gets fixed by all the other steps (mixed precision, code training, scale distill). It didn't get fixed by just using a more chinese imatrix. Thanks for the suggestion! That definitely cleared things up quite a bit.
