You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Gemma-4 E4B-it · AutoRound W4A16 (MLP + attention output)

A 4-bit weight-only compression of google/gemma-4-E4B-it covering the language-model decoder's MLP and attention-output projections, served under vLLM via the gptq_marlin kernel.

Role: fallback candidate for team Godspeed AI's Round 2 entry in the Resilient AI Challenge (image-to-text). The primary submission is g4e4-it-r2-awq-smoke-v0, which extends quantization to the full decoder (q/k/v/o + MLP via AWQ-Marlin) and adds a response-economy chat template. This repo is the conservative weights-only sibling: fewer modules touched, no template modification.

Model information

Property Value
Base model google/gemma-4-E4B-it (weights untouched outside quantized scope)
Quantized scope language_model.layers.*.mlp.{gate,up,down}_proj, *.self_attn.o_proj
Preserved at bf16 q/k/v projections, vision tower, audio tower, embeddings, lm_head, norms
Format W4A16, group size 128, exported GPTQ, served as gptq_marlin
Algorithm AutoRound 0.12.3 MLLM mode, RTN (--iters 0), asymmetric
Size on disk 10.6 GB (vs 16.0 GB bf16, −34%)
Training None — no finetuning, distillation, or healing at any stage

Benchmark results

NVIDIA L4 (evaluator hardware), temperature=1.0, top_p=0.95, top_k=64, max_tokens=180:

Benchmark BF16 base This model Δ energy
9-category image validation (caption, OCR, table, receipt, dashboard, chart, form, handwriting, diagram) 40/41 facts · 0.738 Wh 37/41 facts · ~0.36 Wh −51%
50-row Indic 5×5 (translation, summarization, QA, transliteration, code-switch) 61/65 · 0.557 Wh 59/65 · 0.256 Wh −54%

Reference reproduction on RTX A6000 (Ampere sm_86): −33% (image) and −52% (Indic) at equal or better facts. Per-category recovery stays above 80% in every measured category on both GPUs.

Usage

vllm serve Shankara-A-S/g4e4-it-r2-w4a16-mlpo-v0 --config vllm_config.yaml

The packaged vllm_config.yaml is self-contained (no infrastructure-specific parameters): gpu-memory-utilization 0.90, max-model-len 8192, dtype bfloat16, quantization gptq_marlin, trust-remote-code true.

Implementation notes

  • GPTQ format rejects mixed-precision shards inside vLLM's fused qkv_proj, which is why q/k/v stay bf16 in this artifact. The sister AWQ repo solves that with the AWQ-Marlin pack layout and quantizes the full decoder.
  • vLLM 0.20.2 prints a cosmetic Casting torch.float16 to torch.bfloat16 warning at load; behavior is correct.
  • Gemma 4's heterogeneous head dims (local 256 / global 512) force the TRITON_ATTN backend automatically.
  • Cold start to first /health ≈ 2–3 minutes (weight load + torch.compile); warm restarts ≈ 15 s.

Limitations

  • Image-understanding composition is slightly more sensitive to quantization than document OCR (worst single category ≈ 89% of base).
  • This artifact deliberately leaves ~15% additional energy savings on the table versus the primary submission in exchange for minimal-surface changes.

License

Apache 2.0, inherited from google/gemma-4-E4B-it.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shankara-A-S/g4e4-it-r2-w4a16-mlpo-v0

Quantized
(367)
this model