Qwen3.6-27B-OBLITERATED — GPTQ-Pro Selective 4-bit (g128)

Overview

Qwen3.6-27B-OBLITERATED-GPTQ-Pro-selective-4bit-g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo. It is intended for open-source evaluation, reproducible experimentation, and compatible local or hosted inference workflows. The wording below is deliberately limited to what can be verified from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base OBLITERATUS/Qwen3.6-27B-OBLITERATED
Intended task text-generation
License apache-2.0

What is included

  • *.safetensors (8 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (16 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/Qwen3.6-27B-OBLITERATED-GPTQ-Pro-selective-4bit-g128 \
  --tensor-parallel-size 2 \
  --quantization gptq_marlin \
  --dtype float16 \
  --kv-cache-dtype fp8_e4m3 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

4-bit GPTQ-Pro quantization of OBLITERATUS/Qwen3.6-27B-OBLITERATED, a dense (non-MoE) Qwen3.5-architecture 27B model with a hybrid attention design: 48 of 64 layers use linear/gated attention (GatedDeltaNet-style), the other 16 use regular full GQA attention.

What "selective" means here

This checkpoint quantizes only ~25% of linear layers to 4-bit; the remaining ~75% (all linear-attention internals, all full-attention Q/K/V projections + norms, embeddings, lm_head, and a set of MLP layers identified as sensitive) are kept at full bf16/fp16 precision.

The specific module list was originally built for a selective AWQ quantization recipe — derived from a prior GGUF Q6_K_XL quantization's own precision choices, treated as a proxy for per-module sensitivity. That AWQ attempt turned out to be non-viable for this architecture (see below), so the same preservation list was replayed through GPTQ-Pro's QuantizeConfig.dynamic mechanism instead, which supports skipping arbitrary modules by regex regardless of the model's default quantization plan.

Why not AWQ: AutoRound's AWQ implementation has no registered smooth-balance mapping for this model's hybrid linear-attention block structure and silently falls back to a generic template that doesn't match its actual connectivity. The resulting checkpoint produced degenerate/garbage output (confirmed via three independent inference stacks: vLLM+awq_marlin, vLLM+generic awq kernel, and plain transformers + upstream GPTQModel — all three converged on identical garbage, ruling out a serving/kernel bug and pointing at the quantization stage itself). GPTQ-Pro's algorithm (Hessian-based weight reconstruction, no activation-aware smooth-balance step) has no equivalent architecture-mapping dependency and produces coherent output.

Comparison

Uniform GPTQ-Pro Selective (this repo)
Size 17G 29G
Modules quantized ~all linear layers, all 64 layers ~25%, sensitive modules preserved bf16
Speed (vLLM, TP=2, 3x RTX 3090) ~57 tok/s ~23 tok/s
Output coherent coherent

No MTP/speculative-decoding weights in this checkpoint (verified: 851 total tensors, 0 mtp/vision tensors).

Serving (vLLM)

vllm serve groxaxo/Qwen3.6-27B-OBLITERATED-GPTQ-Pro-selective-4bit-g128 \
  --tensor-parallel-size 2 \
  --quantization gptq_marlin \
  --dtype float16 \
  --kv-cache-dtype fp8_e4m3 \
  --trust-remote-code

Requires ~2x24GB GPUs (or equivalent) given the larger bf16-preserved fraction. gptq_marlin requires dtype float16 (vLLM rejects bfloat16 for this kernel).

Quantization recipe

Produced with GPTQ-Pro (scripts/quant_qwen36_obliterated_gptqpro.py --preset quality --dynamic-ignore-json modules_to_not_convert.generated.json), bits=4, group_size=128, sym=True, 64 calibration samples of real code, seqlen 512.

Downloads last month
15
Safetensors
Model size
27B params
Tensor type
BF16
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for groxaxo/Qwen3.6-27B-OBLITERATED-GPTQ-Pro-selective-4bit-g128

Base model

Qwen/Qwen3.6-27B
Quantized
(5)
this model