Calibration data: sampling strategy, mix ratio and seqlen?

#2
by gxx - opened

Thanks for sharing the detailed GPTQ config! A few follow-ups on the calibration data:

  1. Is the 2,048 samples the total across both datasets, or 2,048 per dataset?
  2. What's the mixing ratio between evol-codealpaca-v1 (code) and C4 (general text)
    — e.g. 1:1, 8:2, or something else?
  3. Were the 2,048 samples drawn directly from the full datasets (random / stratified
    sampling), or did you first curate/filter a smaller pool and then sample 2,048 from it?
  4. What sequence length were the samples truncated / packed to (2048, 4096, ...)?

Trying to reproduce a matching recipe on an internal Qwen3.5-35B-A3B checkpoint. Thanks!

The ratio is not explicitly set in the config I used. It ends up being ~70/30 code/c4, but depends on the max token length selected. I don't truncate anything, I sample the training datasets for complete pairs within the token length bins. If you have the hardware for it sampling at even higher token lengths may be justified. I used 4 bins with the longest token lengths being 2048. This was a hardware limitation for me as I was running into oom issues on anything longer.

Thanks for your reply! Could you please clarify which specific variant of the C4 dataset you used (for example, the en or en.noblocklist variant from allenai/c4)
image

Hi again!

I am currently trying to run the GPTQ quantization (4-bit, group_size 32) for Qwen3.5-35B-A3B on dual A100 (80GB) GPUs, but the process is extremely slow. It takes about 12 minutes per layer, estimating over 8 hours in total.

Here is the command I am running:

python quantize_gptq_qwen35.py \
  --model_path Qwen3.5-35B-A3B \
  --calib_data mixed_evol_codealpaca_c4.json \
  --bits 4 --group_size 32 --calib_samples 2048
![image](https://cdn-uploads.huggingface.co/production/uploads/63da29cd1ab35adf09f298c7/_Pi3T9aCuMcYIWAF8nZO1.png)

I used the en variant.

The timing sounds about right. You are essentially training the quantized model and the moe variants were slow. I think it took me around 13 hours on my setup for 3.5-35B.

As the quantization progressed to Layer 1, the estimated time indeed jumped to over 1 day (currently showing "1 day, 1:47:47" with 1 hour 17 minutes elapsed). It seems the initial 8-hour estimate was a bit optimistic because the self-attention layers processed very quickly, but once it hit the 256 sequential MoE experts, the speed dropped significantly.

I also noticed a few fallback(rtn) warnings in the logs (for example, on mlp.experts.152.gate_proj) due to Hessian instability.

Just to double check:

  1. Did you also experience these fallback(rtn) warnings during your 13-hour run?
  2. Is 24+ hours considered normal for a dual-A100 setup with 2048 calibration samples, or did you use any specific settings to avoid the slow sequential expert loops?

Thank you so much for your patience and insights!
image

Yes rtn fallback is normal for rare experts. I don't know about A100s tbh. I used 4 x Mi100s. It's possible that there are settings which could speed up the quantization on A100s, but I haven't used A100s for quantization.

Sign up or log in to comment