4-bit quant using Intel AutoRound

This is the default W4A16 scheme.

Quantization Details

at time of quantization, the default implied values not listed in the below json are as follows:

{"batch_size": 8, "iters": 200, "seqlen": 2048, "nsamples": 128, "lr": None}

quantization_config.json

{
  "bits": 4,
  "group_size": 128,
  "sym": true,
  "data_type": "int",
  "low_gpu_mem_usage": true,
  "autoround_version": "0.9.2",
  "quant_method": "auto-round",
  "packing_format": "auto_round:auto_gptq"
}
Downloads last month
5
Safetensors
Model size
11B params
Tensor type
I32
BF16
F16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for hoborific/Lilitu-L3.3-70b-0.1-W4A16-AutoRound

Quantized
(3)
this model