dtestnyrr commited on
Commit
79636b3
·
verified ·
1 Parent(s): 4497f1a

Initial GPTQ int4 upload

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: talkie-lm/talkie-1930-13b-it
4
+ tags:
5
+ - gptq
6
+ - 4-bit
7
+ - quantized
8
+ - instruction-tuned
9
+ - vintage-language-model
10
+ - chat
11
+ language:
12
+ - en
13
+ pipeline_tag: text-generation
14
+ ---
15
+
16
+ # talkie-1930-13b-it — GPTQ int4
17
+
18
+ A 4-bit GPTQ quantization of [`talkie-lm/talkie-1930-13b-it`](https://huggingface.co/talkie-lm/talkie-1930-13b-it) — the **instruction-tuned** variant of talkie 13B, by Alec Radford, Nick Levine, and David Duvenaud.
19
+
20
+ The base model was trained on 260B tokens of pre-1931 English. This IT variant was further fine-tuned on a custom instruction-following dataset built entirely from pre-1931 reference works (etiquette manuals, letter-writing manuals, encyclopedias, poetry collections), then refined with online DPO. The result: a chatbot that answers in **early-20th-century formal English** and has no knowledge of anything after 1930.
21
+
22
+ This quantization shrinks the model from ~24.7 GB (bf16) to ~7.4 GB (int4), so it fits comfortably on a single 16 GB consumer GPU.
23
+
24
+ ## Use it
25
+
26
+ ```python
27
+ from gptqmodel import GPTQModel
28
+ import talkie_hf.talkie_qmodel # registers TalkieQModel
29
+
30
+ model = GPTQModel.load("dtestnyrr/talkie-1930-13b-it-gptq-int4", trust_remote_code=True)
31
+
32
+ # Use the talkie chat template
33
+ prompt = "<|user|>Write a brief letter declining a dinner invitation.<|end|><|assistant|>"
34
+ ids = model.tokenizer(prompt, return_tensors="pt").input_ids.cuda()
35
+ out = model.generate(input_ids=ids, max_new_tokens=300, do_sample=True, temperature=0.7)
36
+ print(model.tokenizer.decode(out[0], skip_special_tokens=False))
37
+ ```
38
+
39
+ Stop generation at any of: `<|end|>`, `<|user|>`, `<|assistant|>`, `<|system|>`, `<|endoftext|>`.
40
+
41
+ ## Chat template
42
+
43
+ ```
44
+ <|system|>{optional system prompt}<|end|>
45
+ <|user|>{user message 1}<|end|>
46
+ <|assistant|>{model reply 1}<|end|>
47
+ <|user|>{user message 2}<|end|>
48
+ <|assistant|>
49
+ ```
50
+
51
+ End the prompt with `<|assistant|>` (no trailing `<|end|>`) to signal the model should generate.
52
+
53
+ ## Quantization details
54
+
55
+ Same recipe as the base-variant int4 release ([dtestnyrr/talkie-1930-13b-base-gptq-int4](https://huggingface.co/dtestnyrr/talkie-1930-13b-base-gptq-int4)):
56
+
57
+ | Parameter | Value |
58
+ |---|---|
59
+ | Method | GPTQ |
60
+ | Bits | 4 |
61
+ | Group size | 128 |
62
+ | Activation order | False |
63
+ | Symmetric | True |
64
+ | Effective bits per weight | ~4.29 BPW |
65
+ | Calibration corpus | 256 × 2048-token windows from 8 pre-1931 Project Gutenberg classics |
66
+ | Quantization framework | [GPTQModel](https://github.com/ModelCloud/GPTQModel) v6.0.3 |
67
+
68
+ ## What it's good at
69
+
70
+ - Period-correct prose in any genre (letters, essays, sermons, news editorials, fiction)
71
+ - Encyclopedia-style explanations of anything pre-1930
72
+ - 1900-1930s-style poetry in proper meter and rhyme
73
+ - Edwardian / Georgian etiquette
74
+ - "Predicting" the future from a 1930 vantage point
75
+
76
+ ## What it can't do
77
+
78
+ - Knowledge of anything post-1930 (no WWII, no computers, no internet, no DNA structure, no antibiotics beyond very early research, etc.)
79
+ - Modern slang or casual register — it will respond formally regardless of how casually you address it
80
+ - Technical assistance with modern tools, languages, or frameworks
81
+ - Knowledge of people who became famous after 1930
82
+
83
+ ## Architecture
84
+
85
+ Custom decoder-only transformer (40 layers, 5120 hidden, 40 heads × 128 head_dim, 13,696 SwiGLU intermediate, vocab 65,540 = 65,536 BPE merges + 4 chat special tokens). Uses RoPE θ=10⁶, F.rms_norm everywhere, QK-norm, and per-residual learnable gain modules. See [`modeling_talkie.py`](modeling_talkie.py) for the full HF-compatible port.
86
+
87
+ ## License & attribution
88
+
89
+ Apache 2.0. Original model credit:
90
+ - Authors: Alec Radford, Nick Levine, David Duvenaud
91
+ - Project: https://talkie-lm.com/
92
+ - Reference code: https://github.com/talkie-lm/talkie
config.json ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "TalkieForCausalLM"
4
+ ],
5
+ "auto_map": {
6
+ "AutoConfig": "configuration_talkie.TalkieConfig",
7
+ "AutoModelForCausalLM": "modeling_talkie.TalkieForCausalLM"
8
+ },
9
+ "dtype": "bfloat16",
10
+ "eos_token": "<|endoftext|>",
11
+ "eos_token_id": 262143,
12
+ "head_dim": 128,
13
+ "hidden_size": 5120,
14
+ "intermediate_size": 13696,
15
+ "max_position_embeddings": 2048,
16
+ "model_type": "talkie",
17
+ "num_attention_heads": 40,
18
+ "num_hidden_layers": 40,
19
+ "pad_token_id": 262143,
20
+ "quantization_config": {
21
+ "bits": 4,
22
+ "checkpoint_format": "gptq",
23
+ "desc_act": false,
24
+ "format": "gptq",
25
+ "group_size": 128,
26
+ "lm_head": false,
27
+ "meta": {
28
+ "act_group_aware": true,
29
+ "auto_forward_data_parallel": true,
30
+ "damp_auto_increment": 0.01,
31
+ "damp_percent": 0.05,
32
+ "fallback": {
33
+ "smooth": null,
34
+ "strategy": "rtn",
35
+ "threshold": "0.5%"
36
+ },
37
+ "foem": null,
38
+ "gc_mode": "interval",
39
+ "gptaq": null,
40
+ "hessian": {
41
+ "chunk_bytes": null,
42
+ "chunk_size": null,
43
+ "staging_dtype": "float32"
44
+ },
45
+ "mock_quantization": false,
46
+ "mse": 0.0,
47
+ "offload_to_disk": true,
48
+ "offload_to_disk_path": "./gptqmodel_offload/ctrizrxh-tetrojst/",
49
+ "pack_impl": "cpu",
50
+ "quantizer": [
51
+ "gptqmodel:6.0.3"
52
+ ],
53
+ "static_groups": false,
54
+ "true_sequential": true,
55
+ "uri": "https://github.com/modelcloud/gptqmodel",
56
+ "vram_strategy": "exclusive",
57
+ "wait_for_submodule_finalizers": false
58
+ },
59
+ "method": "gptq",
60
+ "pack_dtype": "int32",
61
+ "quant_method": "gptq",
62
+ "sym": true
63
+ },
64
+ "rope_parameters": {
65
+ "rope_theta": 1000000.0,
66
+ "rope_type": "default"
67
+ },
68
+ "rope_theta": 1000000.0,
69
+ "tie_word_embeddings": false,
70
+ "transformers_version": "5.6.2",
71
+ "use_cache": false,
72
+ "vocab_size": 65540
73
+ }
configuration_talkie.py ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from transformers import PretrainedConfig
2
+
3
+
4
+ class TalkieConfig(PretrainedConfig):
5
+ model_type = "talkie"
6
+
7
+ def __init__(
8
+ self,
9
+ vocab_size: int = 65536,
10
+ hidden_size: int = 5120,
11
+ num_hidden_layers: int = 40,
12
+ num_attention_heads: int = 40,
13
+ head_dim: int = 128,
14
+ intermediate_size: int = 13696,
15
+ max_position_embeddings: int = 2048,
16
+ rope_theta: float = 1_000_000.0,
17
+ tie_word_embeddings: bool = False,
18
+ **kwargs,
19
+ ):
20
+ self.vocab_size = vocab_size
21
+ self.hidden_size = hidden_size
22
+ self.num_hidden_layers = num_hidden_layers
23
+ self.num_attention_heads = num_attention_heads
24
+ self.head_dim = head_dim
25
+ self.intermediate_size = intermediate_size
26
+ self.max_position_embeddings = max_position_embeddings
27
+ self.rope_theta = rope_theta
28
+ super().__init__(tie_word_embeddings=tie_word_embeddings, **kwargs)
generation_config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "do_sample": true,
4
+ "output_attentions": false,
5
+ "output_hidden_states": false,
6
+ "transformers_version": "5.6.2",
7
+ "use_cache": false
8
+ }
model-00001-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d168496e38c3848ee7002557e440c7f57b3469a3d67da11e5a4caaeb649c54b8
3
+ size 4293839946
model-00002-of-00002.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2d77720b8c0dd6c059680823acdc70988afa2324650c8eaabf7735688928a410
3
+ size 3606506944
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
modeling_talkie.py ADDED
@@ -0,0 +1,253 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """HuggingFace-compatible port of TalkieModel.
2
+
3
+ Mirrors talkie/src/talkie/model.py exactly:
4
+ - F.rms_norm everywhere (no learnable RMSNorm scale)
5
+ - RoPE with base=1e6
6
+ - Attention with QK-norm and per-head HeadGain on Q
7
+ - SwiGLU MLP
8
+ - Per-layer ActGain on attn / mlp / embed_skip residuals
9
+ - WeightGain on the lm_head matrix
10
+
11
+ Linear projections are renamed to Llama conventions (q_proj, k_proj, v_proj,
12
+ o_proj, gate_proj, up_proj, down_proj) so GPTQModel auto-detects them.
13
+ """
14
+
15
+ from __future__ import annotations
16
+
17
+ import torch
18
+ import torch.nn as nn
19
+ import torch.nn.functional as F
20
+ from transformers import PreTrainedModel
21
+ from transformers.generation import GenerationMixin
22
+ from transformers.modeling_outputs import CausalLMOutput
23
+
24
+ from .configuration_talkie import TalkieConfig
25
+
26
+
27
+ def _apply_rotary_emb(x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch.Tensor:
28
+ # x: [B, T, H, D]; cos/sin: [1, T, 1, D/2]
29
+ d = x.shape[-1] // 2
30
+ x1, x2 = x[..., :d], x[..., d:]
31
+ y1 = x1 * cos + x2 * sin
32
+ y2 = -x1 * sin + x2 * cos
33
+ return torch.cat([y1, y2], dim=-1).type_as(x)
34
+
35
+
36
+ class HeadGain(nn.Module):
37
+ def __init__(self, n_head: int):
38
+ super().__init__()
39
+ self.head_g = nn.Parameter(torch.ones(n_head))
40
+
41
+ def forward(self, x: torch.Tensor) -> torch.Tensor:
42
+ return x * self.head_g.type_as(x).view(1, 1, -1, 1)
43
+
44
+
45
+ class WeightGain(nn.Module):
46
+ def __init__(self):
47
+ super().__init__()
48
+ self.w_g = nn.Parameter(torch.ones(1))
49
+
50
+ def forward(self, w: torch.Tensor) -> torch.Tensor:
51
+ return w * self.w_g.type_as(w)
52
+
53
+
54
+ class ActGain(nn.Module):
55
+ def __init__(self, init_value: float):
56
+ super().__init__()
57
+ self.a_g = nn.Parameter(torch.ones(1) * init_value)
58
+
59
+ def forward(self, x: torch.Tensor) -> torch.Tensor:
60
+ return x * self.a_g.type_as(x)
61
+
62
+
63
+ class TalkieAttention(nn.Module):
64
+ def __init__(self, config: TalkieConfig):
65
+ super().__init__()
66
+ self.n_head = config.num_attention_heads
67
+ self.head_dim = config.head_dim
68
+ h = config.hidden_size
69
+ self.q_proj = nn.Linear(h, h, bias=False)
70
+ self.k_proj = nn.Linear(h, h, bias=False)
71
+ self.v_proj = nn.Linear(h, h, bias=False)
72
+ self.o_proj = nn.Linear(h, h, bias=False)
73
+ self.head_gain = HeadGain(config.num_attention_heads)
74
+
75
+ def forward(self, x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch.Tensor:
76
+ bsz, seq_len, _ = x.size()
77
+ q = self.q_proj(x).view(bsz, seq_len, self.n_head, self.head_dim)
78
+ k = self.k_proj(x).view(bsz, seq_len, self.n_head, self.head_dim)
79
+ v = self.v_proj(x).view(bsz, seq_len, self.n_head, self.head_dim)
80
+
81
+ q = _apply_rotary_emb(q, cos, sin)
82
+ k = _apply_rotary_emb(k, cos, sin)
83
+ q = F.rms_norm(q, (q.size(-1),))
84
+ k = F.rms_norm(k, (k.size(-1),))
85
+ q = self.head_gain(q)
86
+
87
+ # SDPA expects [B, H, T, D]
88
+ y = F.scaled_dot_product_attention(
89
+ q.transpose(1, 2), k.transpose(1, 2), v.transpose(1, 2), is_causal=True
90
+ )
91
+ y = y.transpose(1, 2).contiguous().view(bsz, seq_len, -1)
92
+ return self.o_proj(y)
93
+
94
+
95
+ class TalkieMLP(nn.Module):
96
+ def __init__(self, config: TalkieConfig):
97
+ super().__init__()
98
+ self.gate_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False)
99
+ self.up_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False)
100
+ self.down_proj = nn.Linear(config.intermediate_size, config.hidden_size, bias=False)
101
+
102
+ def forward(self, x: torch.Tensor) -> torch.Tensor:
103
+ return self.down_proj(F.silu(self.gate_proj(x)) * self.up_proj(x))
104
+
105
+
106
+ class TalkieDecoderLayer(nn.Module):
107
+ def __init__(self, config: TalkieConfig):
108
+ super().__init__()
109
+ self.self_attn = TalkieAttention(config)
110
+ self.attn_gain = ActGain((2 * config.num_hidden_layers) ** -0.5)
111
+ self.mlp = TalkieMLP(config)
112
+ self.mlp_gain = ActGain((2 * config.num_hidden_layers) ** -0.5)
113
+ self.embed_skip = ActGain(0.0)
114
+
115
+ def forward(
116
+ self,
117
+ hidden_states: torch.Tensor,
118
+ e_x: torch.Tensor = None,
119
+ cos: torch.Tensor = None,
120
+ sin: torch.Tensor = None,
121
+ **kwargs, # absorb attention_mask / position_ids / etc that HF tooling injects
122
+ ) -> torch.Tensor:
123
+ x = hidden_states
124
+ x = x + self.attn_gain(self.self_attn(F.rms_norm(x, (x.shape[-1],)), cos, sin))
125
+ x = x + self.mlp_gain(self.mlp(F.rms_norm(x, (x.shape[-1],))))
126
+ x = x + self.embed_skip(e_x)
127
+ return x
128
+
129
+
130
+ class TalkiePreTrainedModel(PreTrainedModel):
131
+ config_class = TalkieConfig
132
+ base_model_prefix = "model"
133
+ supports_gradient_checkpointing = False
134
+ _no_split_modules = ["TalkieDecoderLayer"]
135
+
136
+ def _init_weights(self, module):
137
+ if isinstance(module, nn.Linear):
138
+ module.weight.data.normal_(mean=0.0, std=0.02)
139
+ if module.bias is not None:
140
+ module.bias.data.zero_()
141
+ elif isinstance(module, nn.Embedding):
142
+ module.weight.data.normal_(mean=0.0, std=0.02)
143
+
144
+
145
+ class TalkieModel(TalkiePreTrainedModel):
146
+ def __init__(self, config: TalkieConfig):
147
+ super().__init__(config)
148
+ self.embed_tokens = nn.Embedding(config.vocab_size, config.hidden_size)
149
+ self.layers = nn.ModuleList(
150
+ [TalkieDecoderLayer(config) for _ in range(config.num_hidden_layers)]
151
+ )
152
+ # cos/sin are computed lazily in forward — see _rope. Avoid register_buffer
153
+ # so HF's meta-init / low_cpu_mem_usage loading path does not leave us
154
+ # holding meta tensors that we then try to slice (which raises).
155
+ self._rope_cache: tuple[torch.Tensor, torch.Tensor, torch.device, torch.dtype, int] | None = None
156
+ self.post_init()
157
+
158
+ @staticmethod
159
+ def _build_rope(
160
+ seq_len: int, head_dim: int, base: float, device, dtype
161
+ ) -> tuple[torch.Tensor, torch.Tensor]:
162
+ ch = torch.arange(0, head_dim, 2, dtype=torch.float32, device=device)
163
+ inv_freq = 1.0 / (base ** (ch / head_dim))
164
+ t = torch.arange(seq_len, dtype=torch.float32, device=device)
165
+ freqs = torch.outer(t, inv_freq)
166
+ cos, sin = freqs.cos().to(dtype), freqs.sin().to(dtype)
167
+ return cos[None, :, None, :], sin[None, :, None, :]
168
+
169
+ def _rope(self, seq_len: int, device, dtype) -> tuple[torch.Tensor, torch.Tensor]:
170
+ cache = self._rope_cache
171
+ if (cache is None or cache[2] != device or cache[3] != dtype or cache[4] < seq_len):
172
+ cap = max(seq_len, self.config.max_position_embeddings)
173
+ cos, sin = self._build_rope(cap, self.config.head_dim, self.config.rope_theta, device, dtype)
174
+ self._rope_cache = (cos, sin, device, dtype, cap)
175
+ cos, sin, _, _, _ = self._rope_cache
176
+ return cos[:, :seq_len], sin[:, :seq_len]
177
+
178
+ def forward(self, input_ids: torch.LongTensor, **kwargs) -> torch.Tensor:
179
+ _, seq_len = input_ids.shape
180
+ x = self.embed_tokens(input_ids)
181
+ x = F.rms_norm(x, (x.shape[-1],))
182
+ e_x = x # post-RMSNorm input embeddings; reused as embed_skip source at every layer
183
+ cos, sin = self._rope(seq_len, x.device, x.dtype)
184
+ for layer in self.layers:
185
+ # Pass e_x/cos/sin as kwargs so HF tooling (GPTQModel etc) captures
186
+ # and replays them per-sample when iterating layers individually.
187
+ x = layer(x, e_x=e_x, cos=cos, sin=sin)
188
+ x = F.rms_norm(x, (x.shape[-1],))
189
+ return x
190
+
191
+
192
+ class TalkieForCausalLM(TalkiePreTrainedModel, GenerationMixin):
193
+ _tied_weights_keys = []
194
+ _supports_cache_class = False
195
+ _supports_static_cache = False
196
+ # Talkie has no KV cache implementation — every generate step recomputes
197
+ # the full sequence. Mirror the reference talkie inference behavior.
198
+
199
+ def __init__(self, config: TalkieConfig):
200
+ super().__init__(config)
201
+ # Force use_cache=False so HF generate doesn't try to feed only the
202
+ # last token via past_key_values (which we don't support).
203
+ config.use_cache = False
204
+ self.model = TalkieModel(config)
205
+ self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
206
+ self.lm_head_gain = WeightGain()
207
+ self.post_init()
208
+ # Belt-and-suspenders: also force the generation_config to not cache.
209
+ if hasattr(self, "generation_config") and self.generation_config is not None:
210
+ self.generation_config.use_cache = False
211
+
212
+ def prepare_inputs_for_generation(self, input_ids, **kwargs):
213
+ # Strip past_key_values and always feed the full sequence — talkie
214
+ # has no incremental state.
215
+ kwargs.pop("past_key_values", None)
216
+ kwargs.pop("cache_position", None)
217
+ kwargs["use_cache"] = False
218
+ return {"input_ids": input_ids, **kwargs}
219
+
220
+ def get_input_embeddings(self):
221
+ return self.model.embed_tokens
222
+
223
+ def set_input_embeddings(self, value):
224
+ self.model.embed_tokens = value
225
+
226
+ def get_output_embeddings(self):
227
+ return self.lm_head
228
+
229
+ def set_output_embeddings(self, new_embeddings):
230
+ self.lm_head = new_embeddings
231
+
232
+ def forward(
233
+ self,
234
+ input_ids: torch.LongTensor = None,
235
+ attention_mask: torch.Tensor = None, # accepted but unused (causal-only)
236
+ labels: torch.LongTensor = None,
237
+ **kwargs,
238
+ ) -> CausalLMOutput:
239
+ hidden = self.model(input_ids)
240
+ # WeightGain is scalar-broadcast over the lm_head matrix, so applying
241
+ # it on the linear's output is mathematically identical to pre-scaling
242
+ # the weight (and avoids a cross-module tensor passing pattern that
243
+ # confuses accelerate's device-map hooks).
244
+ logits = self.lm_head(hidden).float() * self.lm_head_gain.w_g.float()
245
+ loss = None
246
+ if labels is not None:
247
+ shift_logits = logits[..., :-1, :].contiguous()
248
+ shift_labels = labels[..., 1:].contiguous()
249
+ loss = F.cross_entropy(
250
+ shift_logits.view(-1, shift_logits.size(-1)),
251
+ shift_labels.view(-1),
252
+ )
253
+ return CausalLMOutput(loss=loss, logits=logits)
quant_log.csv ADDED
@@ -0,0 +1,281 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ layer,module,loss,samples,damp,time
2
+ 0,mlp.gate_proj,0.0003081795,0.05000,17.913
3
+ 0,mlp.up_proj,0.0000542289,0.05000,18.174
4
+ 0,mlp.down_proj,0.0019961544,0.05000,21.241
5
+ 0,self_attn.k_proj,0.0000270423,0.05000,27.808
6
+ 0,self_attn.v_proj,0.0000235055,0.05000,27.895
7
+ 0,self_attn.o_proj,0.0000084061,0.05000,27.915
8
+ 0,self_attn.q_proj,0.0000267042,0.05000,27.942
9
+ 1,mlp.gate_proj,0.0003360091,0.05000,17.789
10
+ 1,mlp.up_proj,0.0000520824,0.05000,17.855
11
+ 1,mlp.down_proj,0.0012718793,0.05000,21.678
12
+ 1,self_attn.k_proj,0.0000274328,0.05000,28.150
13
+ 1,self_attn.q_proj,0.0000271052,0.05000,28.208
14
+ 1,self_attn.v_proj,0.0000236766,0.05000,28.278
15
+ 1,self_attn.o_proj,0.0000074347,0.05000,28.287
16
+ 2,mlp.up_proj,0.0000493462,0.05000,16.887
17
+ 2,mlp.gate_proj,0.0003627762,0.05000,17.242
18
+ 2,mlp.down_proj,0.0006479784,0.05000,20.663
19
+ 2,self_attn.v_proj,0.0000240699,0.05000,27.454
20
+ 2,self_attn.k_proj,0.0000277159,0.05000,27.778
21
+ 2,self_attn.q_proj,0.0000266389,0.05000,27.838
22
+ 2,self_attn.o_proj,0.0000097000,0.05000,27.889
23
+ 3,mlp.up_proj,0.0000229121,0.05000,17.016
24
+ 3,mlp.gate_proj,0.0000504172,0.05000,17.101
25
+ 3,mlp.down_proj,0.0000309956,0.05000,20.434
26
+ 3,self_attn.v_proj,0.0000148267,0.05000,28.005
27
+ 3,self_attn.o_proj,0.0000113674,0.05000,28.217
28
+ 3,self_attn.k_proj,0.0000155096,0.05000,28.254
29
+ 3,self_attn.q_proj,0.0000150627,0.05000,28.260
30
+ 4,mlp.up_proj,0.0000312758,0.05000,16.093
31
+ 4,mlp.gate_proj,0.0000512521,0.05000,17.032
32
+ 4,mlp.down_proj,0.0000542724,0.05000,20.156
33
+ 4,self_attn.v_proj,0.0000170589,0.05000,27.938
34
+ 4,self_attn.q_proj,0.0000153016,0.05000,28.085
35
+ 4,self_attn.k_proj,0.0000155928,0.05000,28.107
36
+ 4,self_attn.o_proj,0.0000104716,0.05000,28.149
37
+ 5,mlp.gate_proj,0.0000609030,0.05000,16.785
38
+ 5,mlp.up_proj,0.0000393444,0.05000,16.951
39
+ 5,mlp.down_proj,0.0000239930,0.05000,20.505
40
+ 5,self_attn.v_proj,0.0000176692,0.05000,28.079
41
+ 5,self_attn.o_proj,0.0000058350,0.05000,28.123
42
+ 5,self_attn.q_proj,0.0000156301,0.05000,28.158
43
+ 5,self_attn.k_proj,0.0000160085,0.05000,28.205
44
+ 6,mlp.gate_proj,0.0000515375,0.05000,17.278
45
+ 6,mlp.up_proj,0.0000411530,0.05000,17.776
46
+ 6,mlp.down_proj,0.0000603061,0.05000,20.828
47
+ 6,self_attn.o_proj,0.0000069052,0.05000,27.807
48
+ 6,self_attn.v_proj,0.0000171747,0.05000,28.234
49
+ 6,self_attn.k_proj,0.0000153144,0.05000,28.444
50
+ 6,self_attn.q_proj,0.0000151436,0.05000,28.471
51
+ 7,mlp.gate_proj,0.0000500004,0.05000,17.776
52
+ 7,mlp.up_proj,0.0000463201,0.05000,17.916
53
+ 7,mlp.down_proj,0.0000528303,0.05000,21.097
54
+ 7,self_attn.o_proj,0.0000056215,0.05000,28.227
55
+ 7,self_attn.q_proj,0.0000153945,0.05000,28.349
56
+ 7,self_attn.v_proj,0.0000174461,0.05000,28.433
57
+ 7,self_attn.k_proj,0.0000159118,0.05000,28.465
58
+ 8,mlp.gate_proj,0.0000503716,0.05000,17.055
59
+ 8,mlp.up_proj,0.0000486011,0.05000,17.603
60
+ 8,mlp.down_proj,0.0000455072,0.05000,21.122
61
+ 8,self_attn.k_proj,0.0000179451,0.05000,28.025
62
+ 8,self_attn.v_proj,0.0000188796,0.05000,28.335
63
+ 8,self_attn.q_proj,0.0000175658,0.05000,28.493
64
+ 8,self_attn.o_proj,0.0000047503,0.05000,28.501
65
+ 9,mlp.up_proj,0.0000556839,0.05000,17.929
66
+ 9,mlp.gate_proj,0.0000614202,0.05000,18.665
67
+ 9,mlp.down_proj,0.0000311427,0.05000,21.902
68
+ 9,self_attn.o_proj,0.0000064314,0.05000,27.886
69
+ 9,self_attn.q_proj,0.0000206849,0.05000,28.458
70
+ 9,self_attn.k_proj,0.0000209098,0.05000,28.482
71
+ 9,self_attn.v_proj,0.0000219568,0.05000,28.514
72
+ 10,mlp.gate_proj,0.0000767580,0.05000,16.931
73
+ 10,mlp.up_proj,0.0000679417,0.05000,17.664
74
+ 10,mlp.down_proj,0.0000403810,0.05000,20.917
75
+ 10,self_attn.k_proj,0.0000250125,0.05000,27.692
76
+ 10,self_attn.v_proj,0.0000258458,0.05000,28.023
77
+ 10,self_attn.o_proj,0.0000067075,0.05000,28.249
78
+ 10,self_attn.q_proj,0.0000249492,0.05000,28.268
79
+ 11,mlp.gate_proj,0.0000920226,0.05000,17.943
80
+ 11,mlp.up_proj,0.0000770476,0.05000,18.249
81
+ 11,mlp.down_proj,0.0000358957,0.05000,21.946
82
+ 11,self_attn.k_proj,0.0000284748,0.05000,27.930
83
+ 11,self_attn.q_proj,0.0000282358,0.05000,28.008
84
+ 11,self_attn.o_proj,0.0000082858,0.05000,28.214
85
+ 11,self_attn.v_proj,0.0000290026,0.05000,28.240
86
+ 12,mlp.up_proj,0.0000864709,0.05000,17.018
87
+ 12,mlp.gate_proj,0.0001036938,0.05000,17.838
88
+ 12,mlp.down_proj,0.0000380611,0.05000,21.033
89
+ 12,self_attn.q_proj,0.0000315321,0.05000,28.575
90
+ 12,self_attn.v_proj,0.0000327975,0.05000,28.766
91
+ 12,self_attn.k_proj,0.0000316397,0.05000,28.826
92
+ 12,self_attn.o_proj,0.0000083556,0.05000,28.853
93
+ 13,mlp.up_proj,0.0000923181,0.05000,17.830
94
+ 13,mlp.gate_proj,0.0001100664,0.05000,17.955
95
+ 13,mlp.down_proj,0.0000405574,0.05000,21.390
96
+ 13,self_attn.k_proj,0.0000350662,0.05000,28.578
97
+ 13,self_attn.v_proj,0.0000353049,0.05000,28.654
98
+ 13,self_attn.o_proj,0.0000075047,0.05000,28.673
99
+ 13,self_attn.q_proj,0.0000347006,0.05000,28.691
100
+ 14,mlp.gate_proj,0.0001040181,0.05000,18.228
101
+ 14,mlp.up_proj,0.0000936042,0.05000,18.411
102
+ 14,mlp.down_proj,0.0355678611,0.05000,22.067
103
+ 14,self_attn.q_proj,0.0000348256,0.05000,28.670
104
+ 14,self_attn.k_proj,0.0000353750,0.05000,28.753
105
+ 14,self_attn.v_proj,0.0000369457,0.05000,28.777
106
+ 14,self_attn.o_proj,0.0000068864,0.05000,28.853
107
+ 15,mlp.up_proj,0.0000964519,0.05000,17.509
108
+ 15,mlp.gate_proj,0.0001058904,0.05000,17.618
109
+ 15,mlp.down_proj,0.0000436096,0.05000,21.055
110
+ 15,self_attn.q_proj,0.0000349336,0.05000,28.645
111
+ 15,self_attn.o_proj,0.0000072430,0.05000,28.764
112
+ 15,self_attn.k_proj,0.0000358868,0.05000,28.789
113
+ 15,self_attn.v_proj,0.0000370624,0.05000,28.815
114
+ 16,mlp.up_proj,0.0001003270,0.05000,17.369
115
+ 16,mlp.gate_proj,0.0001031801,0.05000,17.872
116
+ 16,mlp.down_proj,0.0000429029,0.05000,21.587
117
+ 16,self_attn.k_proj,0.0000367004,0.05000,28.719
118
+ 16,self_attn.o_proj,0.0000060289,0.05000,28.823
119
+ 16,self_attn.v_proj,0.0000392006,0.05000,28.886
120
+ 16,self_attn.q_proj,0.0000359511,0.05000,28.907
121
+ 17,mlp.gate_proj,0.0001028387,0.05000,18.211
122
+ 17,mlp.up_proj,0.0001017818,0.05000,18.231
123
+ 17,mlp.down_proj,0.0000387011,0.05000,21.611
124
+ 17,self_attn.k_proj,0.0000379290,0.05000,28.602
125
+ 17,self_attn.v_proj,0.0000402245,0.05000,29.007
126
+ 17,self_attn.q_proj,0.0000368098,0.05000,29.019
127
+ 17,self_attn.o_proj,0.0000049005,0.05000,29.060
128
+ 18,mlp.gate_proj,0.0001001292,0.05000,17.687
129
+ 18,mlp.up_proj,0.0000996311,0.05000,17.958
130
+ 18,mlp.down_proj,0.0000380027,0.05000,21.696
131
+ 18,self_attn.k_proj,0.0000374850,0.05000,28.734
132
+ 18,self_attn.v_proj,0.0000388674,0.05000,28.870
133
+ 18,self_attn.o_proj,0.0000077347,0.05000,29.026
134
+ 18,self_attn.q_proj,0.0000370627,0.05000,29.042
135
+ 19,mlp.gate_proj,0.0000938521,0.05000,17.762
136
+ 19,mlp.up_proj,0.0000959272,0.05000,18.039
137
+ 19,mlp.down_proj,0.0000342812,0.05000,21.534
138
+ 19,self_attn.o_proj,0.0000068072,0.05000,28.265
139
+ 19,self_attn.q_proj,0.0000348771,0.05000,28.741
140
+ 19,self_attn.k_proj,0.0000353500,0.05000,28.802
141
+ 19,self_attn.v_proj,0.0000368596,0.05000,28.824
142
+ 20,mlp.gate_proj,0.0000895623,0.05000,17.897
143
+ 20,mlp.up_proj,0.0000947136,0.05000,18.148
144
+ 20,mlp.down_proj,0.0000322166,0.05000,21.563
145
+ 20,self_attn.k_proj,0.0000337982,0.05000,28.953
146
+ 20,self_attn.v_proj,0.0000363387,0.05000,29.284
147
+ 20,self_attn.q_proj,0.0000332785,0.05000,29.303
148
+ 20,self_attn.o_proj,0.0000063673,0.05000,29.339
149
+ 21,mlp.gate_proj,0.0000896193,0.05000,17.650
150
+ 21,mlp.up_proj,0.0000960996,0.05000,18.170
151
+ 21,mlp.down_proj,0.0000317455,0.05000,21.673
152
+ 21,self_attn.k_proj,0.0000347393,0.05000,29.186
153
+ 21,self_attn.o_proj,0.0000045270,0.05000,29.212
154
+ 21,self_attn.v_proj,0.0000374702,0.05000,29.263
155
+ 21,self_attn.q_proj,0.0000340738,0.05000,29.362
156
+ 22,mlp.up_proj,0.0000987218,0.05000,17.802
157
+ 22,mlp.gate_proj,0.0000916783,0.05000,17.819
158
+ 22,mlp.down_proj,0.0000354047,0.05000,21.465
159
+ 22,self_attn.v_proj,0.0000384820,0.05000,29.241
160
+ 22,self_attn.k_proj,0.0000358335,0.05000,29.278
161
+ 22,self_attn.o_proj,0.0000055582,0.05000,29.375
162
+ 22,self_attn.q_proj,0.0000351940,0.05000,29.393
163
+ 23,mlp.gate_proj,0.0000988934,0.05000,18.062
164
+ 23,mlp.up_proj,0.0001052049,0.05000,18.496
165
+ 23,mlp.down_proj,0.0000408503,0.05000,21.829
166
+ 23,self_attn.v_proj,0.0000397907,0.05000,29.463
167
+ 23,self_attn.k_proj,0.0000372553,0.05000,29.484
168
+ 23,self_attn.q_proj,0.0000363780,0.05000,29.582
169
+ 23,self_attn.o_proj,0.0000048959,0.05000,29.589
170
+ 24,mlp.gate_proj,0.0001048375,0.05000,18.129
171
+ 24,mlp.up_proj,0.0001094447,0.05000,18.787
172
+ 24,mlp.down_proj,0.0000439513,0.05000,22.425
173
+ 24,self_attn.v_proj,0.0000426328,0.05000,29.455
174
+ 24,self_attn.k_proj,0.0000398516,0.05000,29.500
175
+ 24,self_attn.o_proj,0.0000068015,0.05000,29.571
176
+ 24,self_attn.q_proj,0.0000387764,0.05000,29.642
177
+ 25,mlp.gate_proj,0.0001101101,0.05000,18.471
178
+ 25,mlp.up_proj,0.0001150894,0.05000,18.507
179
+ 25,mlp.down_proj,0.0000502667,0.05000,21.972
180
+ 25,self_attn.k_proj,0.0000414529,0.05000,29.407
181
+ 25,self_attn.q_proj,0.0000402841,0.05000,29.522
182
+ 25,self_attn.v_proj,0.0000460898,0.05000,29.677
183
+ 25,self_attn.o_proj,0.0000049322,0.05000,29.679
184
+ 26,mlp.up_proj,0.0001176010,0.05000,18.294
185
+ 26,mlp.gate_proj,0.0001157366,0.05000,18.715
186
+ 26,mlp.down_proj,0.0000451706,0.05000,22.165
187
+ 26,self_attn.v_proj,0.0000473667,0.05000,29.338
188
+ 26,self_attn.q_proj,0.0000415147,0.05000,29.424
189
+ 26,self_attn.k_proj,0.0000431917,0.05000,29.524
190
+ 26,self_attn.o_proj,0.0000047934,0.05000,29.531
191
+ 27,mlp.up_proj,0.0001159971,0.05000,18.210
192
+ 27,mlp.gate_proj,0.0001153119,0.05000,18.384
193
+ 27,mlp.down_proj,0.0000445333,0.05000,21.814
194
+ 27,self_attn.q_proj,0.0000421407,0.05000,29.062
195
+ 27,self_attn.k_proj,0.0000436215,0.05000,29.213
196
+ 27,self_attn.v_proj,0.0000475841,0.05000,29.256
197
+ 27,self_attn.o_proj,0.0000046830,0.05000,29.273
198
+ 28,mlp.gate_proj,0.0001169334,0.05000,17.645
199
+ 28,mlp.up_proj,0.0001143606,0.05000,18.008
200
+ 28,mlp.down_proj,0.0000337466,0.05000,21.847
201
+ 28,self_attn.o_proj,0.0000032019,0.05000,27.434
202
+ 28,self_attn.k_proj,0.0000432929,0.05000,27.719
203
+ 28,self_attn.v_proj,0.0000462728,0.05000,28.007
204
+ 28,self_attn.q_proj,0.0000419060,0.05000,28.151
205
+ 29,mlp.up_proj,0.0001137857,0.05000,18.228
206
+ 29,mlp.gate_proj,0.0001192500,0.05000,18.240
207
+ 29,mlp.down_proj,0.0000252381,0.05000,21.628
208
+ 29,self_attn.o_proj,0.0000022613,0.05000,28.345
209
+ 29,self_attn.v_proj,0.0000449540,0.05000,28.618
210
+ 29,self_attn.q_proj,0.0000412867,0.05000,28.819
211
+ 29,self_attn.k_proj,0.0000426163,0.05000,28.846
212
+ 30,mlp.up_proj,0.0001078411,0.05000,17.656
213
+ 30,mlp.gate_proj,0.0001135521,0.05000,17.932
214
+ 30,mlp.down_proj,0.0000231547,0.05000,21.266
215
+ 30,self_attn.k_proj,0.0000411614,0.05000,28.448
216
+ 30,self_attn.o_proj,0.0000030266,0.05000,28.505
217
+ 30,self_attn.v_proj,0.0000445966,0.05000,28.653
218
+ 30,self_attn.q_proj,0.0000403358,0.05000,28.710
219
+ 31,mlp.up_proj,0.0001045471,0.05000,17.233
220
+ 31,mlp.gate_proj,0.0001093158,0.05000,17.635
221
+ 31,mlp.down_proj,0.0000235892,0.05000,21.389
222
+ 31,self_attn.k_proj,0.0000396851,0.05000,29.104
223
+ 31,self_attn.o_proj,0.0000026403,0.05000,29.284
224
+ 31,self_attn.q_proj,0.0000388012,0.05000,29.338
225
+ 31,self_attn.v_proj,0.0000428868,0.05000,29.353
226
+ 32,mlp.gate_proj,0.0001087208,0.05000,17.920
227
+ 32,mlp.up_proj,0.0001044456,0.05000,18.058
228
+ 32,mlp.down_proj,0.0000240254,0.05000,21.339
229
+ 32,self_attn.q_proj,0.0000378783,0.05000,28.779
230
+ 32,self_attn.v_proj,0.0000425838,0.05000,28.862
231
+ 32,self_attn.k_proj,0.0000395236,0.05000,29.008
232
+ 32,self_attn.o_proj,0.0000025328,0.05000,29.019
233
+ 33,mlp.gate_proj,0.0001147145,0.05000,17.565
234
+ 33,mlp.up_proj,0.0001067313,0.05000,17.897
235
+ 33,mlp.down_proj,0.0000227866,0.05000,21.379
236
+ 33,self_attn.q_proj,0.0000378068,0.05000,29.419
237
+ 33,self_attn.o_proj,0.0000020157,0.05000,29.626
238
+ 33,self_attn.v_proj,0.0000429422,0.05000,29.716
239
+ 33,self_attn.k_proj,0.0000400440,0.05000,29.738
240
+ 34,mlp.up_proj,0.0000997575,0.05000,17.211
241
+ 34,mlp.gate_proj,0.0001061515,0.05000,18.252
242
+ 34,mlp.down_proj,0.0000202997,0.05000,21.364
243
+ 34,self_attn.v_proj,0.0000405380,0.05000,27.850
244
+ 34,self_attn.q_proj,0.0000365120,0.05000,28.540
245
+ 34,self_attn.o_proj,0.0000024578,0.05000,28.644
246
+ 34,self_attn.k_proj,0.0000380477,0.05000,28.668
247
+ 35,mlp.gate_proj,0.0000984036,0.05000,18.177
248
+ 35,mlp.up_proj,0.0000936725,0.05000,18.248
249
+ 35,mlp.down_proj,0.0000177211,0.05000,21.932
250
+ 35,self_attn.q_proj,0.0000328877,0.05000,27.925
251
+ 35,self_attn.v_proj,0.0000387808,0.05000,28.576
252
+ 35,self_attn.k_proj,0.0000347784,0.05000,28.628
253
+ 35,self_attn.o_proj,0.0000015525,0.05000,28.713
254
+ 36,mlp.gate_proj,0.0000940234,0.05000,18.255
255
+ 36,mlp.up_proj,0.0000899822,0.05000,18.335
256
+ 36,mlp.down_proj,0.0000180931,0.05000,21.839
257
+ 36,self_attn.k_proj,0.0000334674,0.05000,29.467
258
+ 36,self_attn.q_proj,0.0000326385,0.05000,29.651
259
+ 36,self_attn.o_proj,0.0000020520,0.05000,29.665
260
+ 36,self_attn.v_proj,0.0000368466,0.05000,29.685
261
+ 37,mlp.gate_proj,0.0000945316,0.05000,19.052
262
+ 37,mlp.up_proj,0.0000905586,0.05000,19.627
263
+ 37,mlp.down_proj,0.0000239953,0.05000,23.262
264
+ 37,self_attn.k_proj,0.0000336605,0.05000,29.512
265
+ 37,self_attn.o_proj,0.0000015146,0.05000,29.655
266
+ 37,self_attn.q_proj,0.0000327701,0.05000,29.723
267
+ 37,self_attn.v_proj,0.0000377369,0.05000,29.729
268
+ 38,mlp.up_proj,0.0000923737,0.05000,24.743
269
+ 38,mlp.gate_proj,0.0000938228,0.05000,24.822
270
+ 38,mlp.down_proj,0.0000328694,0.05000,28.743
271
+ 38,self_attn.k_proj,0.0000334626,0.05000,29.978
272
+ 38,self_attn.v_proj,0.0000377836,0.05000,30.020
273
+ 38,self_attn.q_proj,0.0000331712,0.05000,30.023
274
+ 38,self_attn.o_proj,0.0000035293,0.05000,30.060
275
+ 39,mlp.gate_proj,0.0000792578,0.05000,24.641
276
+ 39,mlp.up_proj,0.0000821907,0.05000,24.913
277
+ 39,mlp.down_proj,0.0000582872,0.05000,28.858
278
+ 39,self_attn.v_proj,0.0000294953,0.05000,29.453
279
+ 39,self_attn.o_proj,0.0000085167,0.05000,29.570
280
+ 39,self_attn.q_proj,0.0000305774,0.05000,29.662
281
+ 39,self_attn.k_proj,0.0000303001,0.05000,29.671
quantize_config.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bits": 4,
3
+ "group_size": 128,
4
+ "desc_act": false,
5
+ "lm_head": false,
6
+ "method": "gptq",
7
+ "quant_method": "gptq",
8
+ "format": "gptq",
9
+ "checkpoint_format": "gptq",
10
+ "pack_dtype": "int32",
11
+ "meta": {
12
+ "quantizer": [
13
+ "gptqmodel:6.0.3"
14
+ ],
15
+ "uri": "https://github.com/modelcloud/gptqmodel",
16
+ "damp_percent": 0.05,
17
+ "damp_auto_increment": 0.01,
18
+ "static_groups": false,
19
+ "true_sequential": true,
20
+ "mse": 0.0,
21
+ "gptaq": null,
22
+ "foem": null,
23
+ "act_group_aware": true,
24
+ "fallback": {
25
+ "strategy": "rtn",
26
+ "threshold": "0.5%",
27
+ "smooth": null
28
+ },
29
+ "offload_to_disk": true,
30
+ "offload_to_disk_path": "./gptqmodel_offload/ctrizrxh-tetrojst/",
31
+ "pack_impl": "cpu",
32
+ "gc_mode": "interval",
33
+ "wait_for_submodule_finalizers": false,
34
+ "auto_forward_data_parallel": true,
35
+ "vram_strategy": "exclusive",
36
+ "mock_quantization": false,
37
+ "hessian": {
38
+ "chunk_size": null,
39
+ "chunk_bytes": null,
40
+ "staging_dtype": "float32"
41
+ }
42
+ },
43
+ "sym": true
44
+ }
talkie_qmodel.py ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """GPTQModel adapter for the Talkie architecture.
2
+
3
+ Importing this module registers TalkieQModel under model_type='talkie' in
4
+ GPTQModel's MODEL_MAP, so `GPTQModel.load(...)` and `GPTQModel.from_quantized(...)`
5
+ work without manual configuration.
6
+
7
+ Auto-detect produces the same module_tree, so this is purely for the from_quantized
8
+ path (which doesn't run auto-detect — module_tree must be a class attribute).
9
+ """
10
+
11
+ from __future__ import annotations
12
+
13
+ from gptqmodel.models.base import BaseQModel
14
+ from gptqmodel.models.auto import MODEL_MAP, SUPPORTED_MODELS
15
+
16
+
17
+ class TalkieQModel(BaseQModel):
18
+ # talkie uses functional F.rms_norm with no learnable scale, so there's no
19
+ # named pre-lm-head normalization module. Empty string disables that hook.
20
+ pre_lm_head_norm_module = ""
21
+
22
+ # Module tree maps GPTQModel's iteration onto our TalkieDecoderLayer:
23
+ # model.layers.{i}.self_attn.{q_proj,k_proj,v_proj,o_proj}
24
+ # model.layers.{i}.mlp.{gate_proj,up_proj,down_proj}
25
+ # Suffix :0/:1 declares quantization grouping order — q/k/v share input
26
+ # (the post-attn-rmsnorm hidden state), o has a different input (SDPA output).
27
+ # Same for gate/up sharing input vs down. Mirrors LlamaQModel.
28
+ module_tree = [
29
+ "model",
30
+ "layers",
31
+ "#",
32
+ {
33
+ "self_attn": ("q_proj:0", "k_proj:0", "v_proj:0", "o_proj:1"),
34
+ "mlp": ("gate_proj:0", "up_proj:0", "down_proj:1"),
35
+ },
36
+ ]
37
+
38
+
39
+ # Register under model_type='talkie' so GPTQModel.load auto-routes to us.
40
+ MODEL_MAP["talkie"] = TalkieQModel
41
+ if "talkie" not in SUPPORTED_MODELS:
42
+ SUPPORTED_MODELS.append("talkie")
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:250b431ae085034dccfb9a4bf13db6bd1f8c375bb7ed499efd57e7c94ca64a3a
3
+ size 43187818
tokenizer_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": null,
4
+ "clean_up_tokenization_spaces": false,
5
+ "eos_token": "<|endoftext|>",
6
+ "is_local": true,
7
+ "local_files_only": false,
8
+ "model_max_length": 1000000000000000019884624838656,
9
+ "pad_token": "<|endoftext|>",
10
+ "tokenizer_class": "TokenizersBackendFast",
11
+ "unk_token": null,
12
+ "_commit_hash": null
13
+ }