GLM-5.3-Flash selective EXL3 3.0 bpw

This checkpoint uses a custom selective-EXL3 TP=4 layout. It has not passed full-server loading, endpoint, or generated-response acceptance. Vision and MTP execution are also unvalidated. The held-out quality result was measured successfully but fell outside the conservative Q4 control gates described below.

This is a selective 3.0 bpw EXL3 conversion of Z.AI's GLM-5.3-Flash-BF16, pinned to source revision a6c167b62691b2bac901344b65cb651a70f53e43. Compatibility with the later official source revision f12e0fe1f6b2ea274c11a569582edfd99d993c5e was separately checked before release assembly.

Only routed-expert gate/up/down projections in language layers 3–44 are EXL3 K3. Attention, linear attention, indexers, mHC, routers, shared experts, dense layers 0–2, embeddings, LM head, norms, vision, and MTP remain at source BF16.

Status

Claim Status
Source and calibration identity Verified
42 routed layers / 288 experts per layer Encoded and independently verified
Artifact structure, indexes, bytes, and checksums Pass
Real EXL3 kernels on all four TP ranks Pass
CUDA graph capture and replay on all four TP ranks Pass
Held-out BF16 comparison Measured; outside Q4 control gates
Full TP4 collective/server/API generation Not run
Vision and MTP execution Not run

Layout and size

The artifact contains 583,090 indexed tensors and 149.56 GB of files. Routed-expert EXL3 tensors account for 115.49 GB; 2,482 retained source-precision tensors account for 33.84 GB. The format identifier is glm53-selective-exl3-tp4-v1, and a compatible custom loader is required. Stock Transformers and generic EXL3 compatibility are not claimed.

Calibration and coverage

Conversion used 600 sealed rows × 2,048 tokens = 1,228,800 tokens from the pinned ExLlamaV3 calibration bundle plus bounded synthetic coverage rows. Natural top-8 routing covered every expert in every routed layer. Calibration was used only for conversion; this release does not train or fine-tune the base model.

The calibration-row SHA-256 is 1cae9bbcd2beb3879a0c459edfca1fd197043ab204b82189c9361de386d0cae1. The conversion used ExLlamaV3 commit 5f3c537ca9d89893d771256f5c43c93656553fbb.

Validation results

The scoped runtime primitive test loaded real EXL3 gate/up/down tensors for all four TP ranks in two waves on two physical GPUs. Every rank captured and replayed a CUDA graph, with zero graph-replay-versus-eager relative L2 error in that test. This does not prove collectives or complete model serving.

Held-out evaluation used 65,504 next-token positions from 32 contiguous 2,048-token blocks of the Salesforce Wikitext-2 raw test split, disjoint from calibration rows.

Metric BF16 source EXL3 3.0 bpw Conservative gate Result
Cross-entropy 1.16296 1.25186 Measured
Perplexity 3.19940 3.49685 absolute Δ ≤ 5% +9.30%, outside
Forward KL, BF16 → EXL3 0.15251 ≤ 0.15 Outside
Top-1 agreement 87.28% ≥ 80% Pass

The raw reports are under evidence/. The release preserves this result without relabeling it as a quality pass.

Sibling releases

Attribution and license

  • Z.AI: base model and MIT license.
  • Dione: selective EXL3 conversion and validation workflow.
  • ExLlamaV3 / TurboDerp: EXL3 format and conversion implementation.
  • Brandon M. Music: earlier MIT-licensed GLM EXL3/TR3 rank-sliced lineage and the public release-engineering pattern of reproducibility closure, cold KLD receipts, explicit runtime profiles, and content-addressed provenance. See brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw and the separately licensed GLM-5.3 release. No ShapleyMCG code, model weights, calibration corpus, or generated artifact is included or used by this release.
  • Salesforce Research: Wikitext evaluation dataset.

The source LICENSE is included. Use this checkpoint only with a loader that understands its selective GLM-5.3 EXL3 TP=4 contract, and treat full serving, vision, and MTP support as pending until separately demonstrated.

Reproducibility, results, runtime, and provenance

  • reproducibility/ records the public-safe source-closure status and verifier.
  • results/ indexes raw held-out quality receipts without hiding the failed conservative gate.
  • runtime/ separates scoped kernel/CUDA-graph evidence from the still-pending full-server and MTP gates.
  • PROVENANCE.md defines the content-addressed identity chain.
  • RELEASE_STATUS.json is the machine-readable gate ledger for this exact Hub snapshot.
Downloads last month
546
Safetensors
Model size
75B params
Tensor type
F32
·
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0xSero/GLM-5.3-Flash-EXL3-3.0bpw

Quantized
(26)
this model

Dataset used to train 0xSero/GLM-5.3-Flash-EXL3-3.0bpw