cbert33 commited on
Commit
d8fd912
·
verified ·
1 Parent(s): 06d6876

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +103 -3
README.md CHANGED
@@ -7,14 +7,114 @@ library_name: transformers
7
  pipeline_tag: image-text-to-text
8
  tags:
9
  - agnes-ai
 
10
  - reasoning
11
  - multimodal
 
 
 
 
12
  - long-context
13
  - hybrid-attention
14
  ---
15
- <p align="center">
16
- <img width="132" src="assets/agnes_logo.svg" alt="Agnes AI logo">
17
- </p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
  <p align="center">
20
  <a href="https://agnes-ai.com/"><img src="https://img.shields.io/badge/Agnes_AI-Website-3248AF" alt="Agnes AI website"></a>
 
7
  pipeline_tag: image-text-to-text
8
  tags:
9
  - agnes-ai
10
+ - uncensored
11
  - reasoning
12
  - multimodal
13
+ - fp8
14
+ - w8a8
15
+ - vllm
16
+ - sglang
17
  - long-context
18
  - hybrid-attention
19
  ---
20
+
21
+
22
+ # Agnes 3.0 Flash FP8 (W8A8, calibrated)
23
+
24
+ Full FP8 quantization of Agnes-3.0-Flash Preview (33B parameters, multimodal).
25
+ Weights and activations are FP8 e4m3 with static per-tensor scales, calibrated
26
+ with LLM Compressor. This gives you both the model and context at FP8. Text,
27
+ image, and video inputs are preserved, and the chat template is upstream's.
28
+
29
+ The Sglang patch from the original model checkpoint is included here as well.
30
+
31
+ *Note that this model seems to be naturally uncensored. I ran Heretic's assessment
32
+ and it gave 0/300 refusals. So no need to abliterate this model. And with that,
33
+ my usual warning here:*
34
+
35
+ > **Uncensored model:** the language checkpoint has undergone abliteration
36
+ to reduce refusal behavior. Treat outputs as untrusted, apply application-level
37
+ safeguards, and do not assume the model will decline harmful requests.
38
+
39
+ > **User responsibility:** this model is provided without warranty. The
40
+ creators, uploaders, and maintainers are not responsible or liable for what
41
+ others generate, publish, deploy, or otherwise do with this abliterated model.
42
+ Users must operate it responsibly, apply appropriate safeguards, comply with
43
+ applicable law, and respect third-party rights. This model is for research
44
+ purposes only and is not intended for production use.
45
+
46
+
47
+ ## Quantization recipe
48
+
49
+ | Field | Value |
50
+ |---|---|
51
+ | Format | `compressed-tensors`, `float-quantized` |
52
+ | Weights | FP8 (e4m3), per-tensor symmetric, static |
53
+ | Activations | FP8 (e4m3), per-tensor symmetric, static, calibrated |
54
+ | Targets | all `Linear` modules except the protected list below |
55
+ | Toolkit | llm-compressor 0.13.0, compressed-tensors 0.18.0 |
56
+ | Calibration | `HuggingFaceH4/ultrachat_200k`, split `train_sft`, revision `8049631c405ae6576f93f445c6b8166f76f5505a` |
57
+ | Calibration size | 512 samples, max sequence length 2048, batch 1, seed 42 |
58
+
59
+ Kept at original precision: `lm_head`, `embed_tokens`, the full `model.visual`
60
+ tower, the recurrent-state layers of delta attention (`conv1d`, `in_proj_a`,
61
+ `in_proj_b`), and every MTP weight (shipped unquantized in
62
+ `model-mtp.safetensors`).
63
+
64
+ `agnes_quantization_manifest.json` records the full build, including the
65
+ SHA-256 of the calibration prompt ids used, so the run is reproducible piece
66
+ for piece. `recipe.yaml` is the machine-readable form of the table above.
67
+
68
+ ## Why the quantization runs on a fused checkpoint
69
+
70
+ The upstream FFN carries two branches per layer: a main branch (width 17,408)
71
+ and a parallel branch (width 2,048). This checkpoint folds them into a single
72
+ set of projections per layer (`intermediate_size: 19456`,
73
+ `parallel_ffn_intermediate_size: 0`).
74
+
75
+ Fusing in bf16 is concatenation and it is exact: gate/up join along the output
76
+ dimension, down along the input dimension. Static per-tensor FP8 scales cannot
77
+ be merged after the fact, because each branch carries its own scale and a
78
+ single fused matrix has room for exactly one. Quantizing the fused layout
79
+ means every serving tensor is quantized once, from the values the model will
80
+ actually run, and calibration observes the same matrix geometry as inference.
81
+
82
+ ## Numerics
83
+
84
+ The fused layout measures 5.9e-4 full-vocabulary KL against the two-branch
85
+ layout at identical weights. The residual comes from reduction order in the
86
+ fused projections, and it sits below the 6.5e-4 that independent inference
87
+ stacks produce from the same unquantized checkpoint. Perplexity moves 17.05 to
88
+ 17.07 across the fusion.
89
+
90
+ The FP8 conversion itself has no formal before/after eval on this card. Run
91
+ your own benchmarks before trusting it for a workload.
92
+
93
+ ## Files
94
+
95
+ | Group | Contents |
96
+ |---|---|
97
+ | Weights | `model-00001-of-00002.safetensors`, `model-00002-of-00002.safetensors`, `model-mtp.safetensors`, `model.safetensors.index.json` |
98
+ | Config | `config.json`, `generation_config.json`, custom `*_agnes.py` modeling and processor code |
99
+ | Tokenizer | `tokenizer.json`, `tokenizer_config.json`, `chat_template.jinja` |
100
+ | Provenance | `recipe.yaml`, `agnes_quantization_manifest.json` |
101
+ | Extras | `sglang_patch/`, `serve.sh` |
102
+
103
+ The checkpoint is standard `compressed-tensors float-quantized`, so any engine
104
+ that reads that format can load it.
105
+
106
+ ## Upstream and license
107
+
108
+ Derived from the `Agnes-AI/Agnes-3.0-Flash` Preview checkpoint at revision
109
+ `891ce4f9ffb89b22888aa7fcc2bb2f3618867684`, Apache 2.0, unchanged. This is a
110
+ community quantization and is not affiliated with Agnes AI. The open-weight
111
+ Preview differs from the production/API model, and the upstream benchmark
112
+ figures below describe the Preview checkpoint before quantization.
113
+
114
+ ---
115
+
116
+ **The original upstream model card follows, unchanged.**
117
+
118
 
119
  <p align="center">
120
  <a href="https://agnes-ai.com/"><img src="https://img.shields.io/badge/Agnes_AI-Website-3248AF" alt="Agnes AI website"></a>