|
Download README.md from XenonFear128/Agnes-3.0-Flash-32GB-NVFP4-k2: direct link, hf CLI and curl.
- Browser
- Download file 3.16 kB
-
https://huggingface.co/XenonFear128/Agnes-3.0-Flash-32GB-NVFP4-k2/resolve/6dc2b4ec2ba2a29c5ccb18a8e451ac9e2a64a5e9/README.md
- Command line
-
hf download hf://XenonFear128/Agnes-3.0-Flash-32GB-NVFP4-k2@6dc2b4ec2ba2a29c5ccb18a8e451ac9e2a64a5e9/README.md
-
curl -L -o README.md https://huggingface.co/XenonFear128/Agnes-3.0-Flash-32GB-NVFP4-k2/resolve/6dc2b4ec2ba2a29c5ccb18a8e451ac9e2a64a5e9/README.md
3.16 kB
| license: apache-2.0 | |
| base_model: Agnes-AI/Agnes-3.0-Flash | |
| base_model_relation: quantized | |
| library_name: modelopt | |
| pipeline_tag: image-text-to-text | |
| tags: | |
| - quantized | |
| - nvfp4 | |
| - fp8 | |
| - modelopt | |
| - custom_code | |
| - tool-calling | |
| # Agnes-3.0-Flash 32GB NVFP4/k=2 (candidate) | |
| > **Hardware status:** This is a 32GB RTX Blackwell candidate profile. The packaged checkpoint has been validated on an RTX PRO 6000 96GB host; final RTX 5090D memory and throughput acceptance is still required. | |
| This repository contains the mixed NVFP4/FP8 checkpoint intended for a single 32GB Blackwell card with a 262K configured context. | |
| ## Quantization | |
| - FFN main and parallel branches: NVFP4. | |
| - Attention and recurrent QKV/QKVZ projections: FP8. | |
| - Attention/recurrent output projections: NVFP4. | |
| - Embedding, lm_head, MTP, vision linear/position/patch weights: FP8. | |
| - Norms, biases, recurrent state: BF16/FP32. | |
| - Main KV cache: NVFP4. MTP draft KV: FP8 E4M3. | |
| ## Runtime | |
| Use the custom patched SGLang runtime required by this checkpoint. The matching runtime patch archive is [`runtime-scripts-v2.tar.gz`](runtime-scripts-v2.tar.gz) (SHA256 `46d537da5a2ef47921fb85c659492274cc1f3dceaa35499d3dc8e773c89d2a16`). The validated profile is MTP k=2, 144 MiB FlashInfer workspace, max Mamba cache 1, prefill CUDA graph disabled, and ReplaySSM enabled. See [`docs/launch-config.json`](docs/launch-config.json) and [`docs/runtime-policy.json`](docs/runtime-policy.json). | |
| For retrieval-heavy long prompts, run [`docs/prepare_fact_anchor.py`](docs/prepare_fact_anchor.py) before generation. It appends [`docs/fact-anchor-template.txt`](docs/fact-anchor-template.txt)-equivalent tokens at the generation boundary while preserving the configured context limit. In a 260K probe this preserved all three access codes and the 57/18/7/32 inventory relation. | |
| ## Evidence and limits | |
| The combined FP8-vision/k=2/262K run used 261632 input tokens plus 56 output tokens. It produced correct retrieval facts, accepted 38/38 MTP drafts, and measured 106.68 tok/s over the short output. Startup-through-sequential-request NVML peak was 32044 MiB on an RTX PRO 6000 96GB host, leaving 724 MiB relative to a 32GiB budget. The required target gate is at least 500 MiB free; repeat this measurement on the RTX 5090D. | |
| Synthetic visual checks passed 3/3. A strict ten-item quality smoke passed 6/10, mostly due to formatting and reasoning-output expectations. Long-form relation behavior is prompt-sensitive; use the fact anchor for retrieval-heavy tasks. These results do not establish broad capability equivalence or 5090D performance. | |
| See [`docs/validation-k2-262k.json`](docs/validation-k2-262k.json), [`docs/5090D-acceptance-checklist.md`](docs/5090D-acceptance-checklist.md), and [`docs/checkpoint-manifest.json`](docs/checkpoint-manifest.json). | |
| The checkpoint's embedded ModelOpt summary is reproduced in [`docs/quantization-metadata-summary.json`](docs/quantization-metadata-summary.json); KV cache precision is a runtime setting and is therefore not encoded in the weight metadata. | |
| The release metadata/runtime bundle is covered by [`docs/release-manifest.json`](docs/release-manifest.json). | |