coolbho3k's picture
Add lossless page15 Engrams; preserve canonical weights and legacy layout
93a1856 verified
|
Raw
History Blame Contribute Delete
7.81 kB
---
license: mit
base_model: deepseek-ai/DeepSeek-V4.1-Flash
base_model_relation: quantized
pipeline_tag: image-text-to-text
inference: false
tags:
- exl3
- 3bpw
- quantized
- deepseek-v41
- experimental
---
# DeepSeek-V4.1-Flash-EXL3-3bpw
Calibrated **EXL3 3bpw routed-expert quantization** of
[DeepSeek-V4.1-Flash](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/tree/df42c109f1defefcbfcedbe7d905718a12266e40).
The original BF16 vision transformer and aligner are preserved unchanged.
**Upload status:** this repository is being populated. Treat the download as
incomplete until `upload-status.json` reports `complete` and the downloaded
files match `release-manifest.json`.
**Runtime compatibility:** this is an experimental weights release for a
specialized vLLM/EXL3 integration, not a drop-in stock-vLLM or standalone
ExLlamaV3 checkpoint. The reusable public serving bundle is still being
finalized and is not included here yet. Do not assume that a stock
`vllm serve` command can load these weights.
## What is quantized?
- All 40 backbone layers and 15,360 routed experts use calibrated 3-bit
EXL3 MUL1 weights.
- The **3bpw label applies to routed experts**, not every parameter or the
average storage cost of the entire checkpoint. EXL3 metadata and padding
add storage beyond the packed 3-bit payload.
- Vision transformer and aligner retain their original BF16 bytes. Routers,
shared experts and other non-routed tensors are copied unchanged from the
published source, which is itself a mixed-precision/native-FP4 checkpoint.
- Original FP8 engram tables/scales are retained under `engrams/`, for
SSD-backed lookup rather than loading those tables onto the GPUs.
- Original draft/MTP weights are retained under `draft/`. They are not loaded
in the tested recipe; speculative decoding is disabled.
## Measured quality
The full 40-layer candidate was compared with the original native-FP4 model
on 64 held-out records. These are **teacher-forced** measurements, not a
BF16 comparison or a claim about broad downstream task accuracy.
| Text metric | Original source | EXL3 3bpw |
| --- | ---: | ---: |
| Prediction targets | 81,880 | 81,880 |
| Mean negative log likelihood | 1.33438994 | 1.34028314 |
| Perplexity | 3.79767842 | 3.82012498 |
| Correct next-token prediction | 70.71690% | 70.60821% |
Perplexity increases **0.59106%**. Top-choice agreement with the original is
**93.24255%**, which is agreement, not task accuracy. The image-answer slice
contains only 84 prediction targets: it is too small for a broad vision claim.
See `evaluation-summary.json` for aggregate measurements and their scope.
Small candidate-only generated-answer checks produced 123/128 on
ARC-Challenge and 63/64 on GSM8K, with thinking disabled. These are sampled
checks, not full benchmarks or original-model generation comparisons;
calibration overlap and pretraining contamination were not assessed.
The original 24-image question suite scored 18/24. Its six misses remain
failures; separately rephrased diagnostics do not replace them.
## Serving measurements and limitations
The previous working custom runtime served text and images across two DGX
Sparks with TP=2, DCP=2, eager execution and **one active request**. It used
native FP8 KV, original BF16 vision and SSD engrams. FP4 KV is not implemented.
- Short-context decode: approximately **7.18 tokens/s**, with 0.379 s server
time to first token in the measured 24-token-prompt/128-token-output test.
- A synthetic **1,047,990-token** retrieval request recovered all three
planted facts in **4,277.419 s (71.29 minutes)**.
- An identical repeat reused 1,047,808 cached tokens and completed in
**10.070 s**; server time to first token was 2.955 s.
- The two short replies at roughly 1M context decoded at approximately
**2.406 tokens/s**. These are not sustained-throughput benchmarks.
The tested profile used GPU memory utilization 0.90 and a downward-only KV
cap of 1,009,612,800 bytes per GPU. Memory margins are tight and host-specific.
Six concurrent requests or six independent 1M histories are **not qualified**.
Cold 1M prefill remains slow; cached-prefix reuse is substantially faster.
Repeated uncached requests for one image question still varied in the
previous working runtime. Tracing found the first recorded divergence at
DCP attention selected-key ordering; the saved last rows had identical key
membership but different order. A deterministic-order candidate has passed
bounded native-kernel tests, but its full-model repair is not yet verified.
This is a serving-repeatability caveat, not a measurement of quantization loss.
Broader comparative generation, coding and vision evaluation remains open.
Runtime source revisions used in the experiments:
- vLLM: `e47aa780bccf59f59dfa2cbb18e17a10b4fe69ba`
- ExLlamaV3: `6ff3a17ea7f3d0026b273d43239398d57f71b788`
- Base checkpoint: `df42c109f1defefcbfcedbe7d905718a12266e40`
Those revisions alone are insufficient: custom DCP/cache-ownership hooks,
SSD engram integration, streaming loading and bounded GB10 kernels are also
required. No clean external rebuild or stock-runtime compatibility is claimed.
## Files and integrity
| Component | Shards | Tensor payload bytes |
| --- | ---: | ---: |
| Active model | 51 | 215,344,840,128 |
| SSD engram tables/scales | 2 | 202,758,032,400 |
| Retained draft/MTP | 2 | 7,932,874,632 |
Total shard storage including safetensors headers is **426,055,648,464 bytes**.
All 55 shards are byte-identical to the independently verified canonical
candidate. Public metadata is a separate revision; no re-quantization occurred.
Tokenizer, model metadata, source license and historical packaging provenance
are included. Calibration examples, activation captures, operational logs,
credentials and runtime container images are not included.
`release-manifest.json` lists the expected size and SHA256 of every payload.
`verify-release.py` verifies a materialized download using standard Python;
run it before loading a GPU model. Obtain the manifest SHA256 from the completed
upload receipt or another trusted channel, then run:
```bash
hf download coolbho3k/DeepSeek-V4.1-Flash-EXL3-3bpw --local-dir ./DeepSeek-V4.1-Flash-EXL3-3bpw
python3 ./DeepSeek-V4.1-Flash-EXL3-3bpw/verify-release.py --directory ./DeepSeek-V4.1-Flash-EXL3-3bpw --manifest-sha256 MANIFEST_SHA256
```
Integrity verification is not a quality benchmark or proof of publisher
authenticity. Historical records under `provenance/` describe packaging and
the evaluation candidate, not a claim that every runtime issue is resolved.
Selected quantization manifest logical SHA256:
`4639e1cf64c5845bb89ae4f9b3648293014bb8c0c32daee2d19f7df0a3f5e86e`.
Original evaluation-candidate package manifest SHA256:
`7f7cfee4a7dc618196b0699a57034dce3316b83257a7320687d079d72d13ee73`.
The upstream MIT license is retained in `LICENSE`.
## Optional lossless page15 Engram assets
The [two-Spark recipe](https://github.com/coolbho3k/DeepSeek-v4.1-Flash-2x-DGX-Spark) also provides a new packed SSD layout in [engram-page15-v1](engram-page15-v1/README.md). Original `engrams/*.safetensors`, main weights and draft weights are unchanged. Old recipe revisions keep using their pinned original files. New recipes pin the matching reader and packed manifest together; do not replace the original files manually. The layout is lossless, not another quantization.
Cold bulk row retrieval was about 2x faster in component tests; measured whole-model decode changed from 30.35 to 30.78 tok/s and warmed 32K prefill was essentially flat. This is not a claim of a 2x serving speedup or broad quality qualification. Full credit to MiaAI Lab / Wesley Young for the native Engram/cache foundation; the derived reader and integration are AGPL-3.0-only in the recipe.