tacodevs commited on
Commit
b09d28b
·
verified ·
1 Parent(s): 44d24a0

Add files using upload-large-folder tool

Browse files
README.md ADDED
@@ -0,0 +1,162 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: mistral-ai-research-license
4
+ license_link: https://mistral.ai/licenses/MNPL-0.1.md
5
+ language:
6
+ - en
7
+ base_model:
8
+ - tacodevs/Behemoth-T1-123B
9
+ - tacodevs/Behemoth-X-R1-123B
10
+ - TheDrummer/Behemoth-X-123B-v2
11
+ - TheDrummer/Behemoth-R1-123B-v2
12
+ tags:
13
+ - mistral
14
+ - mistral-large
15
+ - 123b
16
+ - roleplay
17
+ - creative-writing
18
+ - thinking
19
+ - reasoning
20
+ - fp8
21
+ - w8a8
22
+ - quantized
23
+ - llm-compressor
24
+ pipeline_tag: text-generation
25
+ library_name: transformers
26
+ ---
27
+
28
+ <div align="center">
29
+
30
+ <img src="https://huggingface.co/tacodevs/Behemoth-T1-123B/resolve/main/images/01_header_main.png" alt="Behemoth-T1" width="100%" />
31
+
32
+ <h1>🌴 Behemoth-T1-123B-FP8 🌴</h1>
33
+
34
+ <p><i>The party where literary craft meets unhinged creative writing — production sweet spot.</i></p>
35
+
36
+ <p>
37
+ <a href="https://huggingface.co/tacodevs/Behemoth-T1-123B"><img src="https://img.shields.io/badge/BF16-tacodevs%2FBehemoth--T1--123B-ff7eb6?style=for-the-badge" alt="BF16"/></a>
38
+ <a href="https://huggingface.co/tacodevs/Behemoth-T1-123B-FP8"><img src="https://img.shields.io/badge/FP8-Behemoth--T1--123B--FP8-7eddff?style=for-the-badge" alt="FP8"/></a>
39
+ <a href="https://huggingface.co/tacodevs/Behemoth-T1-123B-GPTQ"><img src="https://img.shields.io/badge/GPTQ--W4A16-Behemoth--T1--123B--GPTQ-bdff7e?style=for-the-badge" alt="GPTQ"/></a>
40
+ </p>
41
+
42
+ </div>
43
+
44
+ ## ☀️ The pitch
45
+
46
+ This is the **FP8 W8A8 dynamic** quantized version of [tacodevs/Behemoth-T1-123B](https://huggingface.co/tacodevs/Behemoth-T1-123B) — a 123B Mistral Large roleplay model that **thinks like a literary author before it writes like a storyteller**.
47
+
48
+ FP8 is the **production sweet spot**: half the VRAM of BF16, ~99% of the quality (essentially lossless), runs cleanly on Hopper GPUs (H100, H200) with native FP8 acceleration. This is the variant most users should pick.
49
+
50
+ For the full pitch, training details, and the philosophy behind T1, see the [BF16 model card](https://huggingface.co/tacodevs/Behemoth-T1-123B).
51
+
52
+ ## ⚡ This variant
53
+
54
+ | | Value |
55
+ |---|---|
56
+ | **Base** | [tacodevs/Behemoth-T1-123B](https://huggingface.co/tacodevs/Behemoth-T1-123B) (BF16) |
57
+ | **Quantization** | FP8 W8A8 dynamic (8-bit weights, 8-bit activations) |
58
+ | **Calibration** | Data-free (per-tensor scale computed analytically) |
59
+ | **Quantizer** | [llm-compressor](https://github.com/vllm-project/llm-compressor) QuantizationModifier |
60
+ | **Size on disk** | ~115 GB (2× smaller than BF16) |
61
+ | **VRAM (8k ctx)** | **~125 GB → fits on 2× 80 GB or 1× 144 GB GPU** |
62
+ | **Quality vs BF16** | ~99% (essentially lossless) |
63
+ | **Speed** | Faster than BF16 on H100/H200 (native FP8 tensor cores) |
64
+
65
+ ## 🎤 How to use
66
+
67
+ T1 expects a **prefilled `<think>` block** to enter literary thinking mode. Use the same 7 prefill phrases as the BF16 model:
68
+
69
+ ```python
70
+ import openai
71
+
72
+ client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="-")
73
+
74
+ PREFILLS = {
75
+ "analytical": "Ok i need to think about how to respond — what does the character feel right now, what from their experience is relevant, what do they value, and what are they trying to achieve, so",
76
+ "creative": "Ok i need to think as a creative writer — what twist would surprise here? Let me find an engaging new direction nobody saw coming, so",
77
+ "unhinged": "Ok i need to think as an unhinged author — raw, explicit, intense, fully in character with no holding back, so",
78
+ }
79
+
80
+ response = client.chat.completions.create(
81
+ model="tacodevs/Behemoth-T1-123B-FP8",
82
+ messages=[
83
+ {"role": "system", "content": CHARACTER_CARD},
84
+ *conversation_history,
85
+ {"role": "user", "content": user_message},
86
+ {"role": "assistant", "content": f"<think>\n{PREFILLS['creative']}\n"},
87
+ ],
88
+ extra_body={
89
+ "continue_final_message": True,
90
+ "add_generation_prompt": False,
91
+ },
92
+ temperature=0.6,
93
+ max_tokens=2048,
94
+ stop=["[INST]", "</s>"],
95
+ )
96
+ ```
97
+
98
+ ## 🚀 Serving with vLLM
99
+
100
+ ```bash
101
+ vllm serve tacodevs/Behemoth-T1-123B-FP8 \
102
+ --tokenizer-mode auto \
103
+ --tensor-parallel-size 2 \
104
+ --max-model-len 8192 \
105
+ --kv-cache-dtype fp8
106
+ ```
107
+
108
+ **Important**: use `--tokenizer-mode auto`, **not** `mistral` — `mistral_common` mode silently mis-templates merged-LoRA checkpoints.
109
+
110
+ Recommended hardware:
111
+ - **2× H100 80GB** — comfortable, ~62GB per GPU
112
+ - **2× H200 144GB** — luxury, plenty of headroom for 32k+ context
113
+ - **1× H100 NVL 188GB** — single-card option
114
+
115
+ FP8 W8A8 uses **native FP8 tensor cores on Hopper GPUs** for both weights and activations, so this variant is also **faster** than BF16 in practice, not just smaller.
116
+
117
+ ## ✅ Quality notes
118
+
119
+ FP8 W8A8 dynamic quantization is essentially lossless for 100B+ models:
120
+
121
+ - ✅ **Stream-of-consciousness thinking shape** — preserved
122
+ - ✅ **Detail surfacing from character cards** — preserved
123
+ - ✅ **Word-for-word output similarity to BF16** — typically 95%+ token overlap on greedy sampling
124
+ - ✅ **Beats base R1 in side-by-side** — preserved
125
+ - ✅ **Production reliability** — same recipe Mistral officially uses for [Mistral Large 3 FP8](https://docs.vllm.ai/projects/recipes/en/latest/Mistral/Mistral-Large-3.html)
126
+
127
+ If you need the absolute reference quality and have 4× 80 GB GPUs, use the [BF16 reference](https://huggingface.co/tacodevs/Behemoth-T1-123B). For most production use cases, **FP8 is the right pick**.
128
+
129
+ ## 🛠️ Training details (from base T1)
130
+
131
+ T1 is a LoRA distillation of Claude Opus 4.5 literary thinking onto
132
+ [`tacodevs/Behemoth-X-R1-123B`](https://huggingface.co/tacodevs/Behemoth-X-R1-123B)
133
+ (itself an SCE merge of Behemoth-X creative writing + Behemoth-R1 reasoning).
134
+
135
+ | | |
136
+ |---|---|
137
+ | **LoRA rank** | 32 (alpha 64, dropout 0.05, all 7 projection modules) |
138
+ | **Trainable params** | 559M / 123B (0.45%) |
139
+ | **Dataset** | 1000 Claude Opus 4.5 thinking traces on real RP conversations |
140
+ | **Loss masking** | Think-only (only the post-prefill thinking continuation gets loss) |
141
+ | **Sequence length** | 4096 |
142
+ | **Epochs** | 2 |
143
+ | **Final eval loss** | 0.9898 |
144
+
145
+ The LoRA only learns the *shape* of literary thinking. The base model's RP prose engine receives **zero gradient updates** — the underlying creative writing voice is structurally preserved.
146
+
147
+ ## 📜 Citation
148
+
149
+ ```bibtex
150
+ @misc{behemoth-t1-2026,
151
+ title = {Behemoth-T1-123B: Literary Thinking Distillation for RP},
152
+ author = {tacodevs},
153
+ year = {2026},
154
+ url = {https://huggingface.co/tacodevs/Behemoth-T1-123B},
155
+ }
156
+ ```
157
+
158
+ <div align="center">
159
+
160
+ <i>The party doesn't end. We just go to bed.</i>
161
+
162
+ </div>
chat_template.jinja ADDED
@@ -0,0 +1 @@
 
 
1
+ {{ bos_token }}{% for message in messages %}{% if message['role'] == 'user' %}{{ '[INST] ' + message['content'] + '[/INST]' }}{% elif message['role'] == 'system' %}{{ '[SYSTEM_PROMPT] ' + message['content'] + '[/SYSTEM_PROMPT]' }}{% elif message['role'] == 'assistant' %}{{ ' ' + message['content'] + eos_token }}{% else %}{{ raise_exception('Only user, system and assistant roles are supported!') }}{% endif %}{% endfor %}
config.json ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "MistralForCausalLM"
4
+ ],
5
+ "attention_dropout": 0.0,
6
+ "bos_token_id": 1,
7
+ "dtype": "bfloat16",
8
+ "eos_token_id": 2,
9
+ "head_dim": 128,
10
+ "hidden_act": "silu",
11
+ "hidden_size": 12288,
12
+ "initializer_range": 0.02,
13
+ "intermediate_size": 28672,
14
+ "max_position_embeddings": 131072,
15
+ "model_type": "mistral",
16
+ "num_attention_heads": 96,
17
+ "num_hidden_layers": 88,
18
+ "num_key_value_heads": 8,
19
+ "quantization_config": {
20
+ "config_groups": {
21
+ "group_0": {
22
+ "format": "float-quantized",
23
+ "input_activations": {
24
+ "actorder": null,
25
+ "block_structure": null,
26
+ "dynamic": true,
27
+ "group_size": null,
28
+ "num_bits": 8,
29
+ "observer": null,
30
+ "observer_kwargs": {},
31
+ "scale_dtype": null,
32
+ "strategy": "token",
33
+ "symmetric": true,
34
+ "type": "float",
35
+ "zp_dtype": null
36
+ },
37
+ "output_activations": null,
38
+ "targets": [
39
+ "Linear"
40
+ ],
41
+ "weights": {
42
+ "actorder": null,
43
+ "block_structure": null,
44
+ "dynamic": false,
45
+ "group_size": null,
46
+ "num_bits": 8,
47
+ "observer": "memoryless_minmax",
48
+ "observer_kwargs": {},
49
+ "scale_dtype": null,
50
+ "strategy": "channel",
51
+ "symmetric": true,
52
+ "type": "float",
53
+ "zp_dtype": null
54
+ }
55
+ }
56
+ },
57
+ "format": "float-quantized",
58
+ "global_compression_ratio": null,
59
+ "ignore": [
60
+ "lm_head"
61
+ ],
62
+ "kv_cache_scheme": null,
63
+ "quant_method": "compressed-tensors",
64
+ "quantization_status": "compressed",
65
+ "sparsity_config": {},
66
+ "transform_config": {},
67
+ "version": "0.14.0.1"
68
+ },
69
+ "rms_norm_eps": 1e-05,
70
+ "rope_theta": 1000000.0,
71
+ "sliding_window": null,
72
+ "tie_word_embeddings": false,
73
+ "transformers_version": "4.57.6",
74
+ "use_cache": true,
75
+ "vocab_size": 32768
76
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 1,
4
+ "eos_token_id": 2,
5
+ "transformers_version": "4.57.6"
6
+ }
model-00001-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:73a43d695712653aef06313996578a5084544887c188aa3431ff87439e26b93f
3
+ size 4958398128
model-00002-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:be327efcc73492294870c16f13f018b6ff79659be1e4fef10ab35eb70c212115
3
+ size 4832680488
model-00003-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7ddc93d770e854a048546372b62eae484e82f63580620bbeb189f684d4afddb0
3
+ size 4857866304
model-00004-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:877c2b368648de761380e054976a9aee850a6d393a52fd26dcce8c3dda2bb5de
3
+ size 4832680552
model-00005-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b537154cef8ea0e83987820718b9db52e1b09e647770b8b762690ad6159147f0
3
+ size 4857866352
model-00006-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:57c30cdf754255fabbee65205386b9568c046e2b0e076c43896a8db980ba6290
3
+ size 4832680552
model-00007-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0f04df9d2b98877a3204c75d7061427045fc4e80c6ecdcb72c7c1d7a5b85a1b0
3
+ size 4857866352
model-00008-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:27a8648490f1cad1b78d5325887ea70acd11bfd32b2ac1bc0cfee07070fd7717
3
+ size 4832680552
model-00009-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b44e8f5820df9203aebe5f10b2fcd9ae1fc807acc674d0af1885f709de24b26f
3
+ size 4857866352
model-00010-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3e84af8d4e6c7b9762960eb82dfb85f005ecf2ea6eaba0f75623de1a70786440
3
+ size 4832680552
model-00011-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6b3b1317bca753d6deaa36b8bbd5bc5964a5ec9f94c47725cd6baa65284bf0fe
3
+ size 4857866352
model-00012-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e9d9bd77fbe403a656a49bb06dd253a4feffc82431f954671649fa68a7b0be7e
3
+ size 4832680552
model-00013-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b4d544095c778b10502f518b9d0da1e355a67585bef57d860477a6cdfdf4611d
3
+ size 4857866352
model-00014-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9fde53f096768a7ea9cf3a2506ca3e56c2025ae670f09fe8b7ae42fc06e86115
3
+ size 4832680552
model-00015-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:26533fe731c6e12826978f6754e36665a43ca06c092016ee449b2f664b4567e4
3
+ size 4857866352
model-00016-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:80993bac9a9b1c8b9ca2f00b4f3f267ad0cb2f6bad2ef9dfffc9f3d92e87013c
3
+ size 4832680552
model-00017-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:190b8cc158e96390507573e29d3a9b116778fedbba071810f2e5e89d782f5e39
3
+ size 4857866352
model-00018-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83b43b783ef90862bb8e5743a9cf2f48e4932aeb8184e43ee99e630075cdb4c6
3
+ size 4832680552
model-00019-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ae5fce9f51cc34abd4536e751449af8d4fe7b2f4b3df22b81d09765ce52c218b
3
+ size 4857866352
model-00020-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:958266dfc10e5570cb5c70e969cdaae5eb7f99ae85e98b182d4535783c9ac023
3
+ size 4832680552
model-00021-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f9515a89c3dc01f5f12967e71bee23a4402c829c2b47f7258a4d5f334d4fa78b
3
+ size 4857866352
model-00022-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9d1ff4b37b8726da95bf1bbaa4f66cd36ec1e92a01b211d518f6a53858a94d03
3
+ size 4832680552
model-00023-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f7c1901b24f5b936153c2e3b62bf771adf2d40872dac9de74f245f311564aaa7
3
+ size 4857866352
model-00024-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8eef824d2a23280b7a50e51e21c04a2b3d14d223dd4846d514500797e8f3bd62
3
+ size 4832680552
model-00025-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a00bf1cba7614844deca4c718f0a59db60d5c2232ea250b1291a37ff07bfe67f
3
+ size 4857866352
model-00026-of-00026.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1a40641032fe24b4b82552eac056d8b1310fd580f720fe4d42ab808c53695f28
3
+ size 2189695048
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
recipe.yaml ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ QuantizationModifier:
4
+ targets: [Linear]
5
+ ignore: [lm_head]
6
+ scheme: FP8_DYNAMIC
7
+ bypass_divisibility_checks: false
special_tokens_map.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "</s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "unk_token": {
17
+ "content": "<unk>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ }
23
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1b968b8dc352f42192367337c78ccc61e1eaddc6d641a579372d4f20694beb7a
3
+ size 587562
tokenizer_config.json ADDED
The diff for this file is too large to render. See raw diff