Runtime format audit (2026-10-06)
No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This CUDA pack is outside MLX/oMLX/MTPLX export scope.
See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.
AX-DeepSeek-OCR-2-CUDA-AXQ-NVFP4-W4A4
Development preview with native FP4 execution. Native AXQuant RTN converts the original official BF16 checkpoint to NVFP4 W4A4. No AWQ is used. Selected text projections use E2M1 FP4 weights and inputs, E4M3FN block scales for 16 values, and FP32 global scales. Protected vision/projector/separators, routers, embeddings, norms and LM head retain original BF16 payloads.
Source and conversion
- Original source: deepseek-ai/DeepSeek-OCR-2,
immutable revision
aaa02f3811945a91062062994c5c4a3f4c0af2b0(apache-2.0). - Original parameters: 3,389,119,360.
- NVFP4 parameters: 2,602,844,160, across 2,196 tensors.
- Protected tensors: 511, each verified for dtype and value equality.
- Final weights: 3,037,806,416 bytes.
- Corrected AXQuant development commit:
1605cc05538fba8f12a5c8af051e3999ea06e947.
Factory conversion used the explicit numpy-reference weight encoder;
inference was independently tested on both NVIDIA GPUs. This is an unreleased
development build with an OCR adapter; provenance.json records exact code
hashes and original source weight digest. Older published AXQuant wheels do
not provide this new conversion path.
BF16 runtime input capture binds the exact source allocation and calibration image. It observes Linear inputs and replays all source experts on observed hidden states, including unrouted experts, to measure down-projection ranges. The plan records sample counts, maxima and 1.25 activation range headroom. Weight and activation global scales are shared across fused projections and complete expert tables. This is RTN with source-bound range calibration, not measured weight sensitivity or broad quality certification.
Native runtime checks
Official vLLM 0.25.1, PyTorch 2.11.0+cu130, CUDA 13.0:
| GPU | Actual Linear / MoE kernel | Page checks |
|---|---|---|
| GeForce RTX 5090 | CutlassNvFp4LinearKernel / FLASHINFER_CUTLASS |
Passed |
| Thor | CutlassNvFp4LinearKernel / VLLM_CUTLASS |
Passed |
Both runs used identical checkpoint hashes, enabled chunked prefill, recognized
AXQuant NVFP4, Invoice 12345, and Total USD 42.50, and passed the repeated-line
prefix check. Exact outputs, request settings and pinned AMD64/ARM64 images are
in development_runtime_smoke.json. Resident Thor services stayed running.
The same self-generated page was used for calibration and validation. This is not a held-out OCR evaluation. The smoke checks those three text lines and repeated-line prefixes; it does not qualify markup/bounding boxes, general OCR accuracy, concurrency or speed. Layout/prefix tokens may still differ from BF16. Eager-mode/JIT/autotune warnings can affect latency; no speed claim is made. The token minimum of 16 is a smoke request setting, not a general OCR accuracy guarantee.
Reproduce
Use a compatible CUDA vLLM 0.25.1 environment. The included standalone runtime script needs vLLM, PyTorch and Pillow; it does not require installing AXQuant.
hf download AutomatosX/AX-DeepSeek-OCR-2-CUDA-AXQ-NVFP4-W4A4 --local-dir model-nvfp4
python model-nvfp4/examples/ocr_smoke.py \
--model model-nvfp4 --image model-nvfp4/calibration/ocr-smoke-page.png \
--output ocr-smoke.json --require-native-fp4 --memory-fraction 0.30
Use --memory-fraction 0.055 for the tested Thor recipe. Actual documents and
concurrency need independent memory sizing. Native FP4 is explicitly required;
actual quantized kernels and their complete allocation coverage are verified; Marlin or emulation fails. Protected BF16 expert tables may select another supported backend.
The script uses native vLLM OCR support with trust_remote_code=False, BF16
outside selected FP4 operations, eager mode, Torch SDPA vision, Triton attention,
one request, and no-repeat n-gram settings 35/128.
Artifacts
axquant_cuda_plan.json: versioned W4A4 plan, original allocation and calibration.axquant_cuda_manifest.json: core digests and development status.activation_calibration.json,calibration/: exact capture evidence and input page.provenance.json,development_runtime_smoke.json: code/source and runtime evidence.SHA256SUMS.txt: every published payload file digest.examples/ocr_smoke.py: tested native-runtime recipe.LICENSE.txt: byte-preserved original license.
Manifest qualification fields remain runtime_verified=false and
quality_certified=false. Separate evidence describes these exact development
cases; this checkpoint does not claim Tier 1/2 certification or a GA release.
- Downloads last month
- 37
Model tree for AutomatosX/AX-DeepSeek-OCR-2-CUDA-AXQ-NVFP4-W4A4
Base model
deepseek-ai/DeepSeek-OCR-2