GLM-5.3-Flash selective EXL3 3.0 bpw
This checkpoint uses a custom selective-EXL3 TP=4 layout. It has not passed full-server loading, endpoint, or generated-response acceptance. Vision and MTP execution are also unvalidated. The held-out quality result was measured successfully but fell outside the conservative Q4 control gates described below.
This is a selective 3.0 bpw EXL3 conversion of
Z.AI's GLM-5.3-Flash-BF16,
pinned to source revision a6c167b62691b2bac901344b65cb651a70f53e43.
Compatibility with the later official source revision
f12e0fe1f6b2ea274c11a569582edfd99d993c5e was separately checked before
release assembly.
Only routed-expert gate/up/down projections in language layers 3–44 are EXL3 K3. Attention, linear attention, indexers, mHC, routers, shared experts, dense layers 0–2, embeddings, LM head, norms, vision, and MTP remain at source BF16.
Status
| Claim | Status |
|---|---|
| Source and calibration identity | Verified |
| 42 routed layers / 288 experts per layer | Encoded and independently verified |
| Artifact structure, indexes, bytes, and checksums | Pass |
| Real EXL3 kernels on all four TP ranks | Pass |
| CUDA graph capture and replay on all four TP ranks | Pass |
| Held-out BF16 comparison | Measured; outside Q4 control gates |
| Full TP4 collective/server/API generation | Not run |
| Vision and MTP execution | Not run |
Layout and size
The artifact contains 583,090 indexed tensors and 149.56 GB of files.
Routed-expert EXL3 tensors account for 115.49 GB; 2,482 retained
source-precision tensors account for 33.84 GB. The format identifier is
glm53-selective-exl3-tp4-v1, and a compatible custom loader is required.
Stock Transformers and generic EXL3 compatibility are not claimed.
Calibration and coverage
Conversion used 600 sealed rows × 2,048 tokens = 1,228,800 tokens from the pinned ExLlamaV3 calibration bundle plus bounded synthetic coverage rows. Natural top-8 routing covered every expert in every routed layer. Calibration was used only for conversion; this release does not train or fine-tune the base model.
The calibration-row SHA-256 is
1cae9bbcd2beb3879a0c459edfca1fd197043ab204b82189c9361de386d0cae1.
The conversion used ExLlamaV3 commit
5f3c537ca9d89893d771256f5c43c93656553fbb.
Validation results
The scoped runtime primitive test loaded real EXL3 gate/up/down tensors for all four TP ranks in two waves on two physical GPUs. Every rank captured and replayed a CUDA graph, with zero graph-replay-versus-eager relative L2 error in that test. This does not prove collectives or complete model serving.
Held-out evaluation used 65,504 next-token positions from 32 contiguous 2,048-token blocks of the Salesforce Wikitext-2 raw test split, disjoint from calibration rows.
| Metric | BF16 source | EXL3 3.0 bpw | Conservative gate | Result |
|---|---|---|---|---|
| Cross-entropy | 1.16296 | 1.25186 | — | Measured |
| Perplexity | 3.19940 | 3.49685 | absolute Δ ≤ 5% | +9.30%, outside |
| Forward KL, BF16 → EXL3 | — | 0.15251 | ≤ 0.15 | Outside |
| Top-1 agreement | — | 87.28% | ≥ 80% | Pass |
The raw reports are under evidence/. The release preserves this result
without relabeling it as a quality pass.
Sibling releases
- EXL3 suite
- 2.5 bpw — pending
- 2.0 bpw — pending
- 4.0 bpw control
Attribution and license
- Z.AI: base model and MIT license.
- Dione: selective EXL3 conversion and validation workflow.
- ExLlamaV3 / TurboDerp: EXL3 format and conversion implementation.
- Brandon M. Music: earlier MIT-licensed GLM EXL3/TR3 rank-sliced lineage and the public release-engineering pattern of reproducibility closure, cold KLD receipts, explicit runtime profiles, and content-addressed provenance. See brandonmusic/GLM-5.2-EXL3-TR3-3.0bpw and the separately licensed GLM-5.3 release. No ShapleyMCG code, model weights, calibration corpus, or generated artifact is included or used by this release.
- Salesforce Research: Wikitext evaluation dataset.
The source LICENSE is included. Use this checkpoint only with a loader that
understands its selective GLM-5.3 EXL3 TP=4 contract, and treat full serving,
vision, and MTP support as pending until separately demonstrated.
Reproducibility, results, runtime, and provenance
reproducibility/records the public-safe source-closure status and verifier.results/indexes raw held-out quality receipts without hiding the failed conservative gate.runtime/separates scoped kernel/CUDA-graph evidence from the still-pending full-server and MTP gates.PROVENANCE.mddefines the content-addressed identity chain.RELEASE_STATUS.jsonis the machine-readable gate ledger for this exact Hub snapshot.
- Downloads last month
- 546
Model tree for 0xSero/GLM-5.3-Flash-EXL3-3.0bpw
Base model
zai-org/GLM-5.3-Flash-BF16