Text Generation
Safetensors
English
talkie
gptq
4-bit precision
quantized
instruction-tuned
vintage-language-model
chat
custom_code
Instructions to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- vLLM
How to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dtestnyrr/talkie-1930-13b-it-gptq-int4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dtestnyrr/talkie-1930-13b-it-gptq-int4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/dtestnyrr/talkie-1930-13b-it-gptq-int4
- SGLang
How to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dtestnyrr/talkie-1930-13b-it-gptq-int4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dtestnyrr/talkie-1930-13b-it-gptq-int4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dtestnyrr/talkie-1930-13b-it-gptq-int4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dtestnyrr/talkie-1930-13b-it-gptq-int4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with Docker Model Runner:
docker model run hf.co/dtestnyrr/talkie-1930-13b-it-gptq-int4
Initial GPTQ int4 upload
Browse files- .gitattributes +1 -0
- README.md +92 -0
- config.json +73 -0
- configuration_talkie.py +28 -0
- generation_config.json +8 -0
- model-00001-of-00002.safetensors +3 -0
- model-00002-of-00002.safetensors +3 -0
- model.safetensors.index.json +0 -0
- modeling_talkie.py +253 -0
- quant_log.csv +281 -0
- quantize_config.json +44 -0
- talkie_qmodel.py +42 -0
- tokenizer.json +3 -0
- tokenizer_config.json +13 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: talkie-lm/talkie-1930-13b-it
|
| 4 |
+
tags:
|
| 5 |
+
- gptq
|
| 6 |
+
- 4-bit
|
| 7 |
+
- quantized
|
| 8 |
+
- instruction-tuned
|
| 9 |
+
- vintage-language-model
|
| 10 |
+
- chat
|
| 11 |
+
language:
|
| 12 |
+
- en
|
| 13 |
+
pipeline_tag: text-generation
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# talkie-1930-13b-it — GPTQ int4
|
| 17 |
+
|
| 18 |
+
A 4-bit GPTQ quantization of [`talkie-lm/talkie-1930-13b-it`](https://huggingface.co/talkie-lm/talkie-1930-13b-it) — the **instruction-tuned** variant of talkie 13B, by Alec Radford, Nick Levine, and David Duvenaud.
|
| 19 |
+
|
| 20 |
+
The base model was trained on 260B tokens of pre-1931 English. This IT variant was further fine-tuned on a custom instruction-following dataset built entirely from pre-1931 reference works (etiquette manuals, letter-writing manuals, encyclopedias, poetry collections), then refined with online DPO. The result: a chatbot that answers in **early-20th-century formal English** and has no knowledge of anything after 1930.
|
| 21 |
+
|
| 22 |
+
This quantization shrinks the model from ~24.7 GB (bf16) to ~7.4 GB (int4), so it fits comfortably on a single 16 GB consumer GPU.
|
| 23 |
+
|
| 24 |
+
## Use it
|
| 25 |
+
|
| 26 |
+
```python
|
| 27 |
+
from gptqmodel import GPTQModel
|
| 28 |
+
import talkie_hf.talkie_qmodel # registers TalkieQModel
|
| 29 |
+
|
| 30 |
+
model = GPTQModel.load("dtestnyrr/talkie-1930-13b-it-gptq-int4", trust_remote_code=True)
|
| 31 |
+
|
| 32 |
+
# Use the talkie chat template
|
| 33 |
+
prompt = "<|user|>Write a brief letter declining a dinner invitation.<|end|><|assistant|>"
|
| 34 |
+
ids = model.tokenizer(prompt, return_tensors="pt").input_ids.cuda()
|
| 35 |
+
out = model.generate(input_ids=ids, max_new_tokens=300, do_sample=True, temperature=0.7)
|
| 36 |
+
print(model.tokenizer.decode(out[0], skip_special_tokens=False))
|
| 37 |
+
```
|
| 38 |
+
|
| 39 |
+
Stop generation at any of: `<|end|>`, `<|user|>`, `<|assistant|>`, `<|system|>`, `<|endoftext|>`.
|
| 40 |
+
|
| 41 |
+
## Chat template
|
| 42 |
+
|
| 43 |
+
```
|
| 44 |
+
<|system|>{optional system prompt}<|end|>
|
| 45 |
+
<|user|>{user message 1}<|end|>
|
| 46 |
+
<|assistant|>{model reply 1}<|end|>
|
| 47 |
+
<|user|>{user message 2}<|end|>
|
| 48 |
+
<|assistant|>
|
| 49 |
+
```
|
| 50 |
+
|
| 51 |
+
End the prompt with `<|assistant|>` (no trailing `<|end|>`) to signal the model should generate.
|
| 52 |
+
|
| 53 |
+
## Quantization details
|
| 54 |
+
|
| 55 |
+
Same recipe as the base-variant int4 release ([dtestnyrr/talkie-1930-13b-base-gptq-int4](https://huggingface.co/dtestnyrr/talkie-1930-13b-base-gptq-int4)):
|
| 56 |
+
|
| 57 |
+
| Parameter | Value |
|
| 58 |
+
|---|---|
|
| 59 |
+
| Method | GPTQ |
|
| 60 |
+
| Bits | 4 |
|
| 61 |
+
| Group size | 128 |
|
| 62 |
+
| Activation order | False |
|
| 63 |
+
| Symmetric | True |
|
| 64 |
+
| Effective bits per weight | ~4.29 BPW |
|
| 65 |
+
| Calibration corpus | 256 × 2048-token windows from 8 pre-1931 Project Gutenberg classics |
|
| 66 |
+
| Quantization framework | [GPTQModel](https://github.com/ModelCloud/GPTQModel) v6.0.3 |
|
| 67 |
+
|
| 68 |
+
## What it's good at
|
| 69 |
+
|
| 70 |
+
- Period-correct prose in any genre (letters, essays, sermons, news editorials, fiction)
|
| 71 |
+
- Encyclopedia-style explanations of anything pre-1930
|
| 72 |
+
- 1900-1930s-style poetry in proper meter and rhyme
|
| 73 |
+
- Edwardian / Georgian etiquette
|
| 74 |
+
- "Predicting" the future from a 1930 vantage point
|
| 75 |
+
|
| 76 |
+
## What it can't do
|
| 77 |
+
|
| 78 |
+
- Knowledge of anything post-1930 (no WWII, no computers, no internet, no DNA structure, no antibiotics beyond very early research, etc.)
|
| 79 |
+
- Modern slang or casual register — it will respond formally regardless of how casually you address it
|
| 80 |
+
- Technical assistance with modern tools, languages, or frameworks
|
| 81 |
+
- Knowledge of people who became famous after 1930
|
| 82 |
+
|
| 83 |
+
## Architecture
|
| 84 |
+
|
| 85 |
+
Custom decoder-only transformer (40 layers, 5120 hidden, 40 heads × 128 head_dim, 13,696 SwiGLU intermediate, vocab 65,540 = 65,536 BPE merges + 4 chat special tokens). Uses RoPE θ=10⁶, F.rms_norm everywhere, QK-norm, and per-residual learnable gain modules. See [`modeling_talkie.py`](modeling_talkie.py) for the full HF-compatible port.
|
| 86 |
+
|
| 87 |
+
## License & attribution
|
| 88 |
+
|
| 89 |
+
Apache 2.0. Original model credit:
|
| 90 |
+
- Authors: Alec Radford, Nick Levine, David Duvenaud
|
| 91 |
+
- Project: https://talkie-lm.com/
|
| 92 |
+
- Reference code: https://github.com/talkie-lm/talkie
|
config.json
ADDED
|
@@ -0,0 +1,73 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"TalkieForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"auto_map": {
|
| 6 |
+
"AutoConfig": "configuration_talkie.TalkieConfig",
|
| 7 |
+
"AutoModelForCausalLM": "modeling_talkie.TalkieForCausalLM"
|
| 8 |
+
},
|
| 9 |
+
"dtype": "bfloat16",
|
| 10 |
+
"eos_token": "<|endoftext|>",
|
| 11 |
+
"eos_token_id": 262143,
|
| 12 |
+
"head_dim": 128,
|
| 13 |
+
"hidden_size": 5120,
|
| 14 |
+
"intermediate_size": 13696,
|
| 15 |
+
"max_position_embeddings": 2048,
|
| 16 |
+
"model_type": "talkie",
|
| 17 |
+
"num_attention_heads": 40,
|
| 18 |
+
"num_hidden_layers": 40,
|
| 19 |
+
"pad_token_id": 262143,
|
| 20 |
+
"quantization_config": {
|
| 21 |
+
"bits": 4,
|
| 22 |
+
"checkpoint_format": "gptq",
|
| 23 |
+
"desc_act": false,
|
| 24 |
+
"format": "gptq",
|
| 25 |
+
"group_size": 128,
|
| 26 |
+
"lm_head": false,
|
| 27 |
+
"meta": {
|
| 28 |
+
"act_group_aware": true,
|
| 29 |
+
"auto_forward_data_parallel": true,
|
| 30 |
+
"damp_auto_increment": 0.01,
|
| 31 |
+
"damp_percent": 0.05,
|
| 32 |
+
"fallback": {
|
| 33 |
+
"smooth": null,
|
| 34 |
+
"strategy": "rtn",
|
| 35 |
+
"threshold": "0.5%"
|
| 36 |
+
},
|
| 37 |
+
"foem": null,
|
| 38 |
+
"gc_mode": "interval",
|
| 39 |
+
"gptaq": null,
|
| 40 |
+
"hessian": {
|
| 41 |
+
"chunk_bytes": null,
|
| 42 |
+
"chunk_size": null,
|
| 43 |
+
"staging_dtype": "float32"
|
| 44 |
+
},
|
| 45 |
+
"mock_quantization": false,
|
| 46 |
+
"mse": 0.0,
|
| 47 |
+
"offload_to_disk": true,
|
| 48 |
+
"offload_to_disk_path": "./gptqmodel_offload/ctrizrxh-tetrojst/",
|
| 49 |
+
"pack_impl": "cpu",
|
| 50 |
+
"quantizer": [
|
| 51 |
+
"gptqmodel:6.0.3"
|
| 52 |
+
],
|
| 53 |
+
"static_groups": false,
|
| 54 |
+
"true_sequential": true,
|
| 55 |
+
"uri": "https://github.com/modelcloud/gptqmodel",
|
| 56 |
+
"vram_strategy": "exclusive",
|
| 57 |
+
"wait_for_submodule_finalizers": false
|
| 58 |
+
},
|
| 59 |
+
"method": "gptq",
|
| 60 |
+
"pack_dtype": "int32",
|
| 61 |
+
"quant_method": "gptq",
|
| 62 |
+
"sym": true
|
| 63 |
+
},
|
| 64 |
+
"rope_parameters": {
|
| 65 |
+
"rope_theta": 1000000.0,
|
| 66 |
+
"rope_type": "default"
|
| 67 |
+
},
|
| 68 |
+
"rope_theta": 1000000.0,
|
| 69 |
+
"tie_word_embeddings": false,
|
| 70 |
+
"transformers_version": "5.6.2",
|
| 71 |
+
"use_cache": false,
|
| 72 |
+
"vocab_size": 65540
|
| 73 |
+
}
|
configuration_talkie.py
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
from transformers import PretrainedConfig
|
| 2 |
+
|
| 3 |
+
|
| 4 |
+
class TalkieConfig(PretrainedConfig):
|
| 5 |
+
model_type = "talkie"
|
| 6 |
+
|
| 7 |
+
def __init__(
|
| 8 |
+
self,
|
| 9 |
+
vocab_size: int = 65536,
|
| 10 |
+
hidden_size: int = 5120,
|
| 11 |
+
num_hidden_layers: int = 40,
|
| 12 |
+
num_attention_heads: int = 40,
|
| 13 |
+
head_dim: int = 128,
|
| 14 |
+
intermediate_size: int = 13696,
|
| 15 |
+
max_position_embeddings: int = 2048,
|
| 16 |
+
rope_theta: float = 1_000_000.0,
|
| 17 |
+
tie_word_embeddings: bool = False,
|
| 18 |
+
**kwargs,
|
| 19 |
+
):
|
| 20 |
+
self.vocab_size = vocab_size
|
| 21 |
+
self.hidden_size = hidden_size
|
| 22 |
+
self.num_hidden_layers = num_hidden_layers
|
| 23 |
+
self.num_attention_heads = num_attention_heads
|
| 24 |
+
self.head_dim = head_dim
|
| 25 |
+
self.intermediate_size = intermediate_size
|
| 26 |
+
self.max_position_embeddings = max_position_embeddings
|
| 27 |
+
self.rope_theta = rope_theta
|
| 28 |
+
super().__init__(tie_word_embeddings=tie_word_embeddings, **kwargs)
|
generation_config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_from_model_config": true,
|
| 3 |
+
"do_sample": true,
|
| 4 |
+
"output_attentions": false,
|
| 5 |
+
"output_hidden_states": false,
|
| 6 |
+
"transformers_version": "5.6.2",
|
| 7 |
+
"use_cache": false
|
| 8 |
+
}
|
model-00001-of-00002.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d168496e38c3848ee7002557e440c7f57b3469a3d67da11e5a4caaeb649c54b8
|
| 3 |
+
size 4293839946
|
model-00002-of-00002.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:2d77720b8c0dd6c059680823acdc70988afa2324650c8eaabf7735688928a410
|
| 3 |
+
size 3606506944
|
model.safetensors.index.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
modeling_talkie.py
ADDED
|
@@ -0,0 +1,253 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""HuggingFace-compatible port of TalkieModel.
|
| 2 |
+
|
| 3 |
+
Mirrors talkie/src/talkie/model.py exactly:
|
| 4 |
+
- F.rms_norm everywhere (no learnable RMSNorm scale)
|
| 5 |
+
- RoPE with base=1e6
|
| 6 |
+
- Attention with QK-norm and per-head HeadGain on Q
|
| 7 |
+
- SwiGLU MLP
|
| 8 |
+
- Per-layer ActGain on attn / mlp / embed_skip residuals
|
| 9 |
+
- WeightGain on the lm_head matrix
|
| 10 |
+
|
| 11 |
+
Linear projections are renamed to Llama conventions (q_proj, k_proj, v_proj,
|
| 12 |
+
o_proj, gate_proj, up_proj, down_proj) so GPTQModel auto-detects them.
|
| 13 |
+
"""
|
| 14 |
+
|
| 15 |
+
from __future__ import annotations
|
| 16 |
+
|
| 17 |
+
import torch
|
| 18 |
+
import torch.nn as nn
|
| 19 |
+
import torch.nn.functional as F
|
| 20 |
+
from transformers import PreTrainedModel
|
| 21 |
+
from transformers.generation import GenerationMixin
|
| 22 |
+
from transformers.modeling_outputs import CausalLMOutput
|
| 23 |
+
|
| 24 |
+
from .configuration_talkie import TalkieConfig
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
def _apply_rotary_emb(x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch.Tensor:
|
| 28 |
+
# x: [B, T, H, D]; cos/sin: [1, T, 1, D/2]
|
| 29 |
+
d = x.shape[-1] // 2
|
| 30 |
+
x1, x2 = x[..., :d], x[..., d:]
|
| 31 |
+
y1 = x1 * cos + x2 * sin
|
| 32 |
+
y2 = -x1 * sin + x2 * cos
|
| 33 |
+
return torch.cat([y1, y2], dim=-1).type_as(x)
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
class HeadGain(nn.Module):
|
| 37 |
+
def __init__(self, n_head: int):
|
| 38 |
+
super().__init__()
|
| 39 |
+
self.head_g = nn.Parameter(torch.ones(n_head))
|
| 40 |
+
|
| 41 |
+
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
| 42 |
+
return x * self.head_g.type_as(x).view(1, 1, -1, 1)
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
class WeightGain(nn.Module):
|
| 46 |
+
def __init__(self):
|
| 47 |
+
super().__init__()
|
| 48 |
+
self.w_g = nn.Parameter(torch.ones(1))
|
| 49 |
+
|
| 50 |
+
def forward(self, w: torch.Tensor) -> torch.Tensor:
|
| 51 |
+
return w * self.w_g.type_as(w)
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
class ActGain(nn.Module):
|
| 55 |
+
def __init__(self, init_value: float):
|
| 56 |
+
super().__init__()
|
| 57 |
+
self.a_g = nn.Parameter(torch.ones(1) * init_value)
|
| 58 |
+
|
| 59 |
+
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
| 60 |
+
return x * self.a_g.type_as(x)
|
| 61 |
+
|
| 62 |
+
|
| 63 |
+
class TalkieAttention(nn.Module):
|
| 64 |
+
def __init__(self, config: TalkieConfig):
|
| 65 |
+
super().__init__()
|
| 66 |
+
self.n_head = config.num_attention_heads
|
| 67 |
+
self.head_dim = config.head_dim
|
| 68 |
+
h = config.hidden_size
|
| 69 |
+
self.q_proj = nn.Linear(h, h, bias=False)
|
| 70 |
+
self.k_proj = nn.Linear(h, h, bias=False)
|
| 71 |
+
self.v_proj = nn.Linear(h, h, bias=False)
|
| 72 |
+
self.o_proj = nn.Linear(h, h, bias=False)
|
| 73 |
+
self.head_gain = HeadGain(config.num_attention_heads)
|
| 74 |
+
|
| 75 |
+
def forward(self, x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch.Tensor:
|
| 76 |
+
bsz, seq_len, _ = x.size()
|
| 77 |
+
q = self.q_proj(x).view(bsz, seq_len, self.n_head, self.head_dim)
|
| 78 |
+
k = self.k_proj(x).view(bsz, seq_len, self.n_head, self.head_dim)
|
| 79 |
+
v = self.v_proj(x).view(bsz, seq_len, self.n_head, self.head_dim)
|
| 80 |
+
|
| 81 |
+
q = _apply_rotary_emb(q, cos, sin)
|
| 82 |
+
k = _apply_rotary_emb(k, cos, sin)
|
| 83 |
+
q = F.rms_norm(q, (q.size(-1),))
|
| 84 |
+
k = F.rms_norm(k, (k.size(-1),))
|
| 85 |
+
q = self.head_gain(q)
|
| 86 |
+
|
| 87 |
+
# SDPA expects [B, H, T, D]
|
| 88 |
+
y = F.scaled_dot_product_attention(
|
| 89 |
+
q.transpose(1, 2), k.transpose(1, 2), v.transpose(1, 2), is_causal=True
|
| 90 |
+
)
|
| 91 |
+
y = y.transpose(1, 2).contiguous().view(bsz, seq_len, -1)
|
| 92 |
+
return self.o_proj(y)
|
| 93 |
+
|
| 94 |
+
|
| 95 |
+
class TalkieMLP(nn.Module):
|
| 96 |
+
def __init__(self, config: TalkieConfig):
|
| 97 |
+
super().__init__()
|
| 98 |
+
self.gate_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False)
|
| 99 |
+
self.up_proj = nn.Linear(config.hidden_size, config.intermediate_size, bias=False)
|
| 100 |
+
self.down_proj = nn.Linear(config.intermediate_size, config.hidden_size, bias=False)
|
| 101 |
+
|
| 102 |
+
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
| 103 |
+
return self.down_proj(F.silu(self.gate_proj(x)) * self.up_proj(x))
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
class TalkieDecoderLayer(nn.Module):
|
| 107 |
+
def __init__(self, config: TalkieConfig):
|
| 108 |
+
super().__init__()
|
| 109 |
+
self.self_attn = TalkieAttention(config)
|
| 110 |
+
self.attn_gain = ActGain((2 * config.num_hidden_layers) ** -0.5)
|
| 111 |
+
self.mlp = TalkieMLP(config)
|
| 112 |
+
self.mlp_gain = ActGain((2 * config.num_hidden_layers) ** -0.5)
|
| 113 |
+
self.embed_skip = ActGain(0.0)
|
| 114 |
+
|
| 115 |
+
def forward(
|
| 116 |
+
self,
|
| 117 |
+
hidden_states: torch.Tensor,
|
| 118 |
+
e_x: torch.Tensor = None,
|
| 119 |
+
cos: torch.Tensor = None,
|
| 120 |
+
sin: torch.Tensor = None,
|
| 121 |
+
**kwargs, # absorb attention_mask / position_ids / etc that HF tooling injects
|
| 122 |
+
) -> torch.Tensor:
|
| 123 |
+
x = hidden_states
|
| 124 |
+
x = x + self.attn_gain(self.self_attn(F.rms_norm(x, (x.shape[-1],)), cos, sin))
|
| 125 |
+
x = x + self.mlp_gain(self.mlp(F.rms_norm(x, (x.shape[-1],))))
|
| 126 |
+
x = x + self.embed_skip(e_x)
|
| 127 |
+
return x
|
| 128 |
+
|
| 129 |
+
|
| 130 |
+
class TalkiePreTrainedModel(PreTrainedModel):
|
| 131 |
+
config_class = TalkieConfig
|
| 132 |
+
base_model_prefix = "model"
|
| 133 |
+
supports_gradient_checkpointing = False
|
| 134 |
+
_no_split_modules = ["TalkieDecoderLayer"]
|
| 135 |
+
|
| 136 |
+
def _init_weights(self, module):
|
| 137 |
+
if isinstance(module, nn.Linear):
|
| 138 |
+
module.weight.data.normal_(mean=0.0, std=0.02)
|
| 139 |
+
if module.bias is not None:
|
| 140 |
+
module.bias.data.zero_()
|
| 141 |
+
elif isinstance(module, nn.Embedding):
|
| 142 |
+
module.weight.data.normal_(mean=0.0, std=0.02)
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
class TalkieModel(TalkiePreTrainedModel):
|
| 146 |
+
def __init__(self, config: TalkieConfig):
|
| 147 |
+
super().__init__(config)
|
| 148 |
+
self.embed_tokens = nn.Embedding(config.vocab_size, config.hidden_size)
|
| 149 |
+
self.layers = nn.ModuleList(
|
| 150 |
+
[TalkieDecoderLayer(config) for _ in range(config.num_hidden_layers)]
|
| 151 |
+
)
|
| 152 |
+
# cos/sin are computed lazily in forward — see _rope. Avoid register_buffer
|
| 153 |
+
# so HF's meta-init / low_cpu_mem_usage loading path does not leave us
|
| 154 |
+
# holding meta tensors that we then try to slice (which raises).
|
| 155 |
+
self._rope_cache: tuple[torch.Tensor, torch.Tensor, torch.device, torch.dtype, int] | None = None
|
| 156 |
+
self.post_init()
|
| 157 |
+
|
| 158 |
+
@staticmethod
|
| 159 |
+
def _build_rope(
|
| 160 |
+
seq_len: int, head_dim: int, base: float, device, dtype
|
| 161 |
+
) -> tuple[torch.Tensor, torch.Tensor]:
|
| 162 |
+
ch = torch.arange(0, head_dim, 2, dtype=torch.float32, device=device)
|
| 163 |
+
inv_freq = 1.0 / (base ** (ch / head_dim))
|
| 164 |
+
t = torch.arange(seq_len, dtype=torch.float32, device=device)
|
| 165 |
+
freqs = torch.outer(t, inv_freq)
|
| 166 |
+
cos, sin = freqs.cos().to(dtype), freqs.sin().to(dtype)
|
| 167 |
+
return cos[None, :, None, :], sin[None, :, None, :]
|
| 168 |
+
|
| 169 |
+
def _rope(self, seq_len: int, device, dtype) -> tuple[torch.Tensor, torch.Tensor]:
|
| 170 |
+
cache = self._rope_cache
|
| 171 |
+
if (cache is None or cache[2] != device or cache[3] != dtype or cache[4] < seq_len):
|
| 172 |
+
cap = max(seq_len, self.config.max_position_embeddings)
|
| 173 |
+
cos, sin = self._build_rope(cap, self.config.head_dim, self.config.rope_theta, device, dtype)
|
| 174 |
+
self._rope_cache = (cos, sin, device, dtype, cap)
|
| 175 |
+
cos, sin, _, _, _ = self._rope_cache
|
| 176 |
+
return cos[:, :seq_len], sin[:, :seq_len]
|
| 177 |
+
|
| 178 |
+
def forward(self, input_ids: torch.LongTensor, **kwargs) -> torch.Tensor:
|
| 179 |
+
_, seq_len = input_ids.shape
|
| 180 |
+
x = self.embed_tokens(input_ids)
|
| 181 |
+
x = F.rms_norm(x, (x.shape[-1],))
|
| 182 |
+
e_x = x # post-RMSNorm input embeddings; reused as embed_skip source at every layer
|
| 183 |
+
cos, sin = self._rope(seq_len, x.device, x.dtype)
|
| 184 |
+
for layer in self.layers:
|
| 185 |
+
# Pass e_x/cos/sin as kwargs so HF tooling (GPTQModel etc) captures
|
| 186 |
+
# and replays them per-sample when iterating layers individually.
|
| 187 |
+
x = layer(x, e_x=e_x, cos=cos, sin=sin)
|
| 188 |
+
x = F.rms_norm(x, (x.shape[-1],))
|
| 189 |
+
return x
|
| 190 |
+
|
| 191 |
+
|
| 192 |
+
class TalkieForCausalLM(TalkiePreTrainedModel, GenerationMixin):
|
| 193 |
+
_tied_weights_keys = []
|
| 194 |
+
_supports_cache_class = False
|
| 195 |
+
_supports_static_cache = False
|
| 196 |
+
# Talkie has no KV cache implementation — every generate step recomputes
|
| 197 |
+
# the full sequence. Mirror the reference talkie inference behavior.
|
| 198 |
+
|
| 199 |
+
def __init__(self, config: TalkieConfig):
|
| 200 |
+
super().__init__(config)
|
| 201 |
+
# Force use_cache=False so HF generate doesn't try to feed only the
|
| 202 |
+
# last token via past_key_values (which we don't support).
|
| 203 |
+
config.use_cache = False
|
| 204 |
+
self.model = TalkieModel(config)
|
| 205 |
+
self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
|
| 206 |
+
self.lm_head_gain = WeightGain()
|
| 207 |
+
self.post_init()
|
| 208 |
+
# Belt-and-suspenders: also force the generation_config to not cache.
|
| 209 |
+
if hasattr(self, "generation_config") and self.generation_config is not None:
|
| 210 |
+
self.generation_config.use_cache = False
|
| 211 |
+
|
| 212 |
+
def prepare_inputs_for_generation(self, input_ids, **kwargs):
|
| 213 |
+
# Strip past_key_values and always feed the full sequence — talkie
|
| 214 |
+
# has no incremental state.
|
| 215 |
+
kwargs.pop("past_key_values", None)
|
| 216 |
+
kwargs.pop("cache_position", None)
|
| 217 |
+
kwargs["use_cache"] = False
|
| 218 |
+
return {"input_ids": input_ids, **kwargs}
|
| 219 |
+
|
| 220 |
+
def get_input_embeddings(self):
|
| 221 |
+
return self.model.embed_tokens
|
| 222 |
+
|
| 223 |
+
def set_input_embeddings(self, value):
|
| 224 |
+
self.model.embed_tokens = value
|
| 225 |
+
|
| 226 |
+
def get_output_embeddings(self):
|
| 227 |
+
return self.lm_head
|
| 228 |
+
|
| 229 |
+
def set_output_embeddings(self, new_embeddings):
|
| 230 |
+
self.lm_head = new_embeddings
|
| 231 |
+
|
| 232 |
+
def forward(
|
| 233 |
+
self,
|
| 234 |
+
input_ids: torch.LongTensor = None,
|
| 235 |
+
attention_mask: torch.Tensor = None, # accepted but unused (causal-only)
|
| 236 |
+
labels: torch.LongTensor = None,
|
| 237 |
+
**kwargs,
|
| 238 |
+
) -> CausalLMOutput:
|
| 239 |
+
hidden = self.model(input_ids)
|
| 240 |
+
# WeightGain is scalar-broadcast over the lm_head matrix, so applying
|
| 241 |
+
# it on the linear's output is mathematically identical to pre-scaling
|
| 242 |
+
# the weight (and avoids a cross-module tensor passing pattern that
|
| 243 |
+
# confuses accelerate's device-map hooks).
|
| 244 |
+
logits = self.lm_head(hidden).float() * self.lm_head_gain.w_g.float()
|
| 245 |
+
loss = None
|
| 246 |
+
if labels is not None:
|
| 247 |
+
shift_logits = logits[..., :-1, :].contiguous()
|
| 248 |
+
shift_labels = labels[..., 1:].contiguous()
|
| 249 |
+
loss = F.cross_entropy(
|
| 250 |
+
shift_logits.view(-1, shift_logits.size(-1)),
|
| 251 |
+
shift_labels.view(-1),
|
| 252 |
+
)
|
| 253 |
+
return CausalLMOutput(loss=loss, logits=logits)
|
quant_log.csv
ADDED
|
@@ -0,0 +1,281 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
layer,module,loss,samples,damp,time
|
| 2 |
+
0,mlp.gate_proj,0.0003081795,0.05000,17.913
|
| 3 |
+
0,mlp.up_proj,0.0000542289,0.05000,18.174
|
| 4 |
+
0,mlp.down_proj,0.0019961544,0.05000,21.241
|
| 5 |
+
0,self_attn.k_proj,0.0000270423,0.05000,27.808
|
| 6 |
+
0,self_attn.v_proj,0.0000235055,0.05000,27.895
|
| 7 |
+
0,self_attn.o_proj,0.0000084061,0.05000,27.915
|
| 8 |
+
0,self_attn.q_proj,0.0000267042,0.05000,27.942
|
| 9 |
+
1,mlp.gate_proj,0.0003360091,0.05000,17.789
|
| 10 |
+
1,mlp.up_proj,0.0000520824,0.05000,17.855
|
| 11 |
+
1,mlp.down_proj,0.0012718793,0.05000,21.678
|
| 12 |
+
1,self_attn.k_proj,0.0000274328,0.05000,28.150
|
| 13 |
+
1,self_attn.q_proj,0.0000271052,0.05000,28.208
|
| 14 |
+
1,self_attn.v_proj,0.0000236766,0.05000,28.278
|
| 15 |
+
1,self_attn.o_proj,0.0000074347,0.05000,28.287
|
| 16 |
+
2,mlp.up_proj,0.0000493462,0.05000,16.887
|
| 17 |
+
2,mlp.gate_proj,0.0003627762,0.05000,17.242
|
| 18 |
+
2,mlp.down_proj,0.0006479784,0.05000,20.663
|
| 19 |
+
2,self_attn.v_proj,0.0000240699,0.05000,27.454
|
| 20 |
+
2,self_attn.k_proj,0.0000277159,0.05000,27.778
|
| 21 |
+
2,self_attn.q_proj,0.0000266389,0.05000,27.838
|
| 22 |
+
2,self_attn.o_proj,0.0000097000,0.05000,27.889
|
| 23 |
+
3,mlp.up_proj,0.0000229121,0.05000,17.016
|
| 24 |
+
3,mlp.gate_proj,0.0000504172,0.05000,17.101
|
| 25 |
+
3,mlp.down_proj,0.0000309956,0.05000,20.434
|
| 26 |
+
3,self_attn.v_proj,0.0000148267,0.05000,28.005
|
| 27 |
+
3,self_attn.o_proj,0.0000113674,0.05000,28.217
|
| 28 |
+
3,self_attn.k_proj,0.0000155096,0.05000,28.254
|
| 29 |
+
3,self_attn.q_proj,0.0000150627,0.05000,28.260
|
| 30 |
+
4,mlp.up_proj,0.0000312758,0.05000,16.093
|
| 31 |
+
4,mlp.gate_proj,0.0000512521,0.05000,17.032
|
| 32 |
+
4,mlp.down_proj,0.0000542724,0.05000,20.156
|
| 33 |
+
4,self_attn.v_proj,0.0000170589,0.05000,27.938
|
| 34 |
+
4,self_attn.q_proj,0.0000153016,0.05000,28.085
|
| 35 |
+
4,self_attn.k_proj,0.0000155928,0.05000,28.107
|
| 36 |
+
4,self_attn.o_proj,0.0000104716,0.05000,28.149
|
| 37 |
+
5,mlp.gate_proj,0.0000609030,0.05000,16.785
|
| 38 |
+
5,mlp.up_proj,0.0000393444,0.05000,16.951
|
| 39 |
+
5,mlp.down_proj,0.0000239930,0.05000,20.505
|
| 40 |
+
5,self_attn.v_proj,0.0000176692,0.05000,28.079
|
| 41 |
+
5,self_attn.o_proj,0.0000058350,0.05000,28.123
|
| 42 |
+
5,self_attn.q_proj,0.0000156301,0.05000,28.158
|
| 43 |
+
5,self_attn.k_proj,0.0000160085,0.05000,28.205
|
| 44 |
+
6,mlp.gate_proj,0.0000515375,0.05000,17.278
|
| 45 |
+
6,mlp.up_proj,0.0000411530,0.05000,17.776
|
| 46 |
+
6,mlp.down_proj,0.0000603061,0.05000,20.828
|
| 47 |
+
6,self_attn.o_proj,0.0000069052,0.05000,27.807
|
| 48 |
+
6,self_attn.v_proj,0.0000171747,0.05000,28.234
|
| 49 |
+
6,self_attn.k_proj,0.0000153144,0.05000,28.444
|
| 50 |
+
6,self_attn.q_proj,0.0000151436,0.05000,28.471
|
| 51 |
+
7,mlp.gate_proj,0.0000500004,0.05000,17.776
|
| 52 |
+
7,mlp.up_proj,0.0000463201,0.05000,17.916
|
| 53 |
+
7,mlp.down_proj,0.0000528303,0.05000,21.097
|
| 54 |
+
7,self_attn.o_proj,0.0000056215,0.05000,28.227
|
| 55 |
+
7,self_attn.q_proj,0.0000153945,0.05000,28.349
|
| 56 |
+
7,self_attn.v_proj,0.0000174461,0.05000,28.433
|
| 57 |
+
7,self_attn.k_proj,0.0000159118,0.05000,28.465
|
| 58 |
+
8,mlp.gate_proj,0.0000503716,0.05000,17.055
|
| 59 |
+
8,mlp.up_proj,0.0000486011,0.05000,17.603
|
| 60 |
+
8,mlp.down_proj,0.0000455072,0.05000,21.122
|
| 61 |
+
8,self_attn.k_proj,0.0000179451,0.05000,28.025
|
| 62 |
+
8,self_attn.v_proj,0.0000188796,0.05000,28.335
|
| 63 |
+
8,self_attn.q_proj,0.0000175658,0.05000,28.493
|
| 64 |
+
8,self_attn.o_proj,0.0000047503,0.05000,28.501
|
| 65 |
+
9,mlp.up_proj,0.0000556839,0.05000,17.929
|
| 66 |
+
9,mlp.gate_proj,0.0000614202,0.05000,18.665
|
| 67 |
+
9,mlp.down_proj,0.0000311427,0.05000,21.902
|
| 68 |
+
9,self_attn.o_proj,0.0000064314,0.05000,27.886
|
| 69 |
+
9,self_attn.q_proj,0.0000206849,0.05000,28.458
|
| 70 |
+
9,self_attn.k_proj,0.0000209098,0.05000,28.482
|
| 71 |
+
9,self_attn.v_proj,0.0000219568,0.05000,28.514
|
| 72 |
+
10,mlp.gate_proj,0.0000767580,0.05000,16.931
|
| 73 |
+
10,mlp.up_proj,0.0000679417,0.05000,17.664
|
| 74 |
+
10,mlp.down_proj,0.0000403810,0.05000,20.917
|
| 75 |
+
10,self_attn.k_proj,0.0000250125,0.05000,27.692
|
| 76 |
+
10,self_attn.v_proj,0.0000258458,0.05000,28.023
|
| 77 |
+
10,self_attn.o_proj,0.0000067075,0.05000,28.249
|
| 78 |
+
10,self_attn.q_proj,0.0000249492,0.05000,28.268
|
| 79 |
+
11,mlp.gate_proj,0.0000920226,0.05000,17.943
|
| 80 |
+
11,mlp.up_proj,0.0000770476,0.05000,18.249
|
| 81 |
+
11,mlp.down_proj,0.0000358957,0.05000,21.946
|
| 82 |
+
11,self_attn.k_proj,0.0000284748,0.05000,27.930
|
| 83 |
+
11,self_attn.q_proj,0.0000282358,0.05000,28.008
|
| 84 |
+
11,self_attn.o_proj,0.0000082858,0.05000,28.214
|
| 85 |
+
11,self_attn.v_proj,0.0000290026,0.05000,28.240
|
| 86 |
+
12,mlp.up_proj,0.0000864709,0.05000,17.018
|
| 87 |
+
12,mlp.gate_proj,0.0001036938,0.05000,17.838
|
| 88 |
+
12,mlp.down_proj,0.0000380611,0.05000,21.033
|
| 89 |
+
12,self_attn.q_proj,0.0000315321,0.05000,28.575
|
| 90 |
+
12,self_attn.v_proj,0.0000327975,0.05000,28.766
|
| 91 |
+
12,self_attn.k_proj,0.0000316397,0.05000,28.826
|
| 92 |
+
12,self_attn.o_proj,0.0000083556,0.05000,28.853
|
| 93 |
+
13,mlp.up_proj,0.0000923181,0.05000,17.830
|
| 94 |
+
13,mlp.gate_proj,0.0001100664,0.05000,17.955
|
| 95 |
+
13,mlp.down_proj,0.0000405574,0.05000,21.390
|
| 96 |
+
13,self_attn.k_proj,0.0000350662,0.05000,28.578
|
| 97 |
+
13,self_attn.v_proj,0.0000353049,0.05000,28.654
|
| 98 |
+
13,self_attn.o_proj,0.0000075047,0.05000,28.673
|
| 99 |
+
13,self_attn.q_proj,0.0000347006,0.05000,28.691
|
| 100 |
+
14,mlp.gate_proj,0.0001040181,0.05000,18.228
|
| 101 |
+
14,mlp.up_proj,0.0000936042,0.05000,18.411
|
| 102 |
+
14,mlp.down_proj,0.0355678611,0.05000,22.067
|
| 103 |
+
14,self_attn.q_proj,0.0000348256,0.05000,28.670
|
| 104 |
+
14,self_attn.k_proj,0.0000353750,0.05000,28.753
|
| 105 |
+
14,self_attn.v_proj,0.0000369457,0.05000,28.777
|
| 106 |
+
14,self_attn.o_proj,0.0000068864,0.05000,28.853
|
| 107 |
+
15,mlp.up_proj,0.0000964519,0.05000,17.509
|
| 108 |
+
15,mlp.gate_proj,0.0001058904,0.05000,17.618
|
| 109 |
+
15,mlp.down_proj,0.0000436096,0.05000,21.055
|
| 110 |
+
15,self_attn.q_proj,0.0000349336,0.05000,28.645
|
| 111 |
+
15,self_attn.o_proj,0.0000072430,0.05000,28.764
|
| 112 |
+
15,self_attn.k_proj,0.0000358868,0.05000,28.789
|
| 113 |
+
15,self_attn.v_proj,0.0000370624,0.05000,28.815
|
| 114 |
+
16,mlp.up_proj,0.0001003270,0.05000,17.369
|
| 115 |
+
16,mlp.gate_proj,0.0001031801,0.05000,17.872
|
| 116 |
+
16,mlp.down_proj,0.0000429029,0.05000,21.587
|
| 117 |
+
16,self_attn.k_proj,0.0000367004,0.05000,28.719
|
| 118 |
+
16,self_attn.o_proj,0.0000060289,0.05000,28.823
|
| 119 |
+
16,self_attn.v_proj,0.0000392006,0.05000,28.886
|
| 120 |
+
16,self_attn.q_proj,0.0000359511,0.05000,28.907
|
| 121 |
+
17,mlp.gate_proj,0.0001028387,0.05000,18.211
|
| 122 |
+
17,mlp.up_proj,0.0001017818,0.05000,18.231
|
| 123 |
+
17,mlp.down_proj,0.0000387011,0.05000,21.611
|
| 124 |
+
17,self_attn.k_proj,0.0000379290,0.05000,28.602
|
| 125 |
+
17,self_attn.v_proj,0.0000402245,0.05000,29.007
|
| 126 |
+
17,self_attn.q_proj,0.0000368098,0.05000,29.019
|
| 127 |
+
17,self_attn.o_proj,0.0000049005,0.05000,29.060
|
| 128 |
+
18,mlp.gate_proj,0.0001001292,0.05000,17.687
|
| 129 |
+
18,mlp.up_proj,0.0000996311,0.05000,17.958
|
| 130 |
+
18,mlp.down_proj,0.0000380027,0.05000,21.696
|
| 131 |
+
18,self_attn.k_proj,0.0000374850,0.05000,28.734
|
| 132 |
+
18,self_attn.v_proj,0.0000388674,0.05000,28.870
|
| 133 |
+
18,self_attn.o_proj,0.0000077347,0.05000,29.026
|
| 134 |
+
18,self_attn.q_proj,0.0000370627,0.05000,29.042
|
| 135 |
+
19,mlp.gate_proj,0.0000938521,0.05000,17.762
|
| 136 |
+
19,mlp.up_proj,0.0000959272,0.05000,18.039
|
| 137 |
+
19,mlp.down_proj,0.0000342812,0.05000,21.534
|
| 138 |
+
19,self_attn.o_proj,0.0000068072,0.05000,28.265
|
| 139 |
+
19,self_attn.q_proj,0.0000348771,0.05000,28.741
|
| 140 |
+
19,self_attn.k_proj,0.0000353500,0.05000,28.802
|
| 141 |
+
19,self_attn.v_proj,0.0000368596,0.05000,28.824
|
| 142 |
+
20,mlp.gate_proj,0.0000895623,0.05000,17.897
|
| 143 |
+
20,mlp.up_proj,0.0000947136,0.05000,18.148
|
| 144 |
+
20,mlp.down_proj,0.0000322166,0.05000,21.563
|
| 145 |
+
20,self_attn.k_proj,0.0000337982,0.05000,28.953
|
| 146 |
+
20,self_attn.v_proj,0.0000363387,0.05000,29.284
|
| 147 |
+
20,self_attn.q_proj,0.0000332785,0.05000,29.303
|
| 148 |
+
20,self_attn.o_proj,0.0000063673,0.05000,29.339
|
| 149 |
+
21,mlp.gate_proj,0.0000896193,0.05000,17.650
|
| 150 |
+
21,mlp.up_proj,0.0000960996,0.05000,18.170
|
| 151 |
+
21,mlp.down_proj,0.0000317455,0.05000,21.673
|
| 152 |
+
21,self_attn.k_proj,0.0000347393,0.05000,29.186
|
| 153 |
+
21,self_attn.o_proj,0.0000045270,0.05000,29.212
|
| 154 |
+
21,self_attn.v_proj,0.0000374702,0.05000,29.263
|
| 155 |
+
21,self_attn.q_proj,0.0000340738,0.05000,29.362
|
| 156 |
+
22,mlp.up_proj,0.0000987218,0.05000,17.802
|
| 157 |
+
22,mlp.gate_proj,0.0000916783,0.05000,17.819
|
| 158 |
+
22,mlp.down_proj,0.0000354047,0.05000,21.465
|
| 159 |
+
22,self_attn.v_proj,0.0000384820,0.05000,29.241
|
| 160 |
+
22,self_attn.k_proj,0.0000358335,0.05000,29.278
|
| 161 |
+
22,self_attn.o_proj,0.0000055582,0.05000,29.375
|
| 162 |
+
22,self_attn.q_proj,0.0000351940,0.05000,29.393
|
| 163 |
+
23,mlp.gate_proj,0.0000988934,0.05000,18.062
|
| 164 |
+
23,mlp.up_proj,0.0001052049,0.05000,18.496
|
| 165 |
+
23,mlp.down_proj,0.0000408503,0.05000,21.829
|
| 166 |
+
23,self_attn.v_proj,0.0000397907,0.05000,29.463
|
| 167 |
+
23,self_attn.k_proj,0.0000372553,0.05000,29.484
|
| 168 |
+
23,self_attn.q_proj,0.0000363780,0.05000,29.582
|
| 169 |
+
23,self_attn.o_proj,0.0000048959,0.05000,29.589
|
| 170 |
+
24,mlp.gate_proj,0.0001048375,0.05000,18.129
|
| 171 |
+
24,mlp.up_proj,0.0001094447,0.05000,18.787
|
| 172 |
+
24,mlp.down_proj,0.0000439513,0.05000,22.425
|
| 173 |
+
24,self_attn.v_proj,0.0000426328,0.05000,29.455
|
| 174 |
+
24,self_attn.k_proj,0.0000398516,0.05000,29.500
|
| 175 |
+
24,self_attn.o_proj,0.0000068015,0.05000,29.571
|
| 176 |
+
24,self_attn.q_proj,0.0000387764,0.05000,29.642
|
| 177 |
+
25,mlp.gate_proj,0.0001101101,0.05000,18.471
|
| 178 |
+
25,mlp.up_proj,0.0001150894,0.05000,18.507
|
| 179 |
+
25,mlp.down_proj,0.0000502667,0.05000,21.972
|
| 180 |
+
25,self_attn.k_proj,0.0000414529,0.05000,29.407
|
| 181 |
+
25,self_attn.q_proj,0.0000402841,0.05000,29.522
|
| 182 |
+
25,self_attn.v_proj,0.0000460898,0.05000,29.677
|
| 183 |
+
25,self_attn.o_proj,0.0000049322,0.05000,29.679
|
| 184 |
+
26,mlp.up_proj,0.0001176010,0.05000,18.294
|
| 185 |
+
26,mlp.gate_proj,0.0001157366,0.05000,18.715
|
| 186 |
+
26,mlp.down_proj,0.0000451706,0.05000,22.165
|
| 187 |
+
26,self_attn.v_proj,0.0000473667,0.05000,29.338
|
| 188 |
+
26,self_attn.q_proj,0.0000415147,0.05000,29.424
|
| 189 |
+
26,self_attn.k_proj,0.0000431917,0.05000,29.524
|
| 190 |
+
26,self_attn.o_proj,0.0000047934,0.05000,29.531
|
| 191 |
+
27,mlp.up_proj,0.0001159971,0.05000,18.210
|
| 192 |
+
27,mlp.gate_proj,0.0001153119,0.05000,18.384
|
| 193 |
+
27,mlp.down_proj,0.0000445333,0.05000,21.814
|
| 194 |
+
27,self_attn.q_proj,0.0000421407,0.05000,29.062
|
| 195 |
+
27,self_attn.k_proj,0.0000436215,0.05000,29.213
|
| 196 |
+
27,self_attn.v_proj,0.0000475841,0.05000,29.256
|
| 197 |
+
27,self_attn.o_proj,0.0000046830,0.05000,29.273
|
| 198 |
+
28,mlp.gate_proj,0.0001169334,0.05000,17.645
|
| 199 |
+
28,mlp.up_proj,0.0001143606,0.05000,18.008
|
| 200 |
+
28,mlp.down_proj,0.0000337466,0.05000,21.847
|
| 201 |
+
28,self_attn.o_proj,0.0000032019,0.05000,27.434
|
| 202 |
+
28,self_attn.k_proj,0.0000432929,0.05000,27.719
|
| 203 |
+
28,self_attn.v_proj,0.0000462728,0.05000,28.007
|
| 204 |
+
28,self_attn.q_proj,0.0000419060,0.05000,28.151
|
| 205 |
+
29,mlp.up_proj,0.0001137857,0.05000,18.228
|
| 206 |
+
29,mlp.gate_proj,0.0001192500,0.05000,18.240
|
| 207 |
+
29,mlp.down_proj,0.0000252381,0.05000,21.628
|
| 208 |
+
29,self_attn.o_proj,0.0000022613,0.05000,28.345
|
| 209 |
+
29,self_attn.v_proj,0.0000449540,0.05000,28.618
|
| 210 |
+
29,self_attn.q_proj,0.0000412867,0.05000,28.819
|
| 211 |
+
29,self_attn.k_proj,0.0000426163,0.05000,28.846
|
| 212 |
+
30,mlp.up_proj,0.0001078411,0.05000,17.656
|
| 213 |
+
30,mlp.gate_proj,0.0001135521,0.05000,17.932
|
| 214 |
+
30,mlp.down_proj,0.0000231547,0.05000,21.266
|
| 215 |
+
30,self_attn.k_proj,0.0000411614,0.05000,28.448
|
| 216 |
+
30,self_attn.o_proj,0.0000030266,0.05000,28.505
|
| 217 |
+
30,self_attn.v_proj,0.0000445966,0.05000,28.653
|
| 218 |
+
30,self_attn.q_proj,0.0000403358,0.05000,28.710
|
| 219 |
+
31,mlp.up_proj,0.0001045471,0.05000,17.233
|
| 220 |
+
31,mlp.gate_proj,0.0001093158,0.05000,17.635
|
| 221 |
+
31,mlp.down_proj,0.0000235892,0.05000,21.389
|
| 222 |
+
31,self_attn.k_proj,0.0000396851,0.05000,29.104
|
| 223 |
+
31,self_attn.o_proj,0.0000026403,0.05000,29.284
|
| 224 |
+
31,self_attn.q_proj,0.0000388012,0.05000,29.338
|
| 225 |
+
31,self_attn.v_proj,0.0000428868,0.05000,29.353
|
| 226 |
+
32,mlp.gate_proj,0.0001087208,0.05000,17.920
|
| 227 |
+
32,mlp.up_proj,0.0001044456,0.05000,18.058
|
| 228 |
+
32,mlp.down_proj,0.0000240254,0.05000,21.339
|
| 229 |
+
32,self_attn.q_proj,0.0000378783,0.05000,28.779
|
| 230 |
+
32,self_attn.v_proj,0.0000425838,0.05000,28.862
|
| 231 |
+
32,self_attn.k_proj,0.0000395236,0.05000,29.008
|
| 232 |
+
32,self_attn.o_proj,0.0000025328,0.05000,29.019
|
| 233 |
+
33,mlp.gate_proj,0.0001147145,0.05000,17.565
|
| 234 |
+
33,mlp.up_proj,0.0001067313,0.05000,17.897
|
| 235 |
+
33,mlp.down_proj,0.0000227866,0.05000,21.379
|
| 236 |
+
33,self_attn.q_proj,0.0000378068,0.05000,29.419
|
| 237 |
+
33,self_attn.o_proj,0.0000020157,0.05000,29.626
|
| 238 |
+
33,self_attn.v_proj,0.0000429422,0.05000,29.716
|
| 239 |
+
33,self_attn.k_proj,0.0000400440,0.05000,29.738
|
| 240 |
+
34,mlp.up_proj,0.0000997575,0.05000,17.211
|
| 241 |
+
34,mlp.gate_proj,0.0001061515,0.05000,18.252
|
| 242 |
+
34,mlp.down_proj,0.0000202997,0.05000,21.364
|
| 243 |
+
34,self_attn.v_proj,0.0000405380,0.05000,27.850
|
| 244 |
+
34,self_attn.q_proj,0.0000365120,0.05000,28.540
|
| 245 |
+
34,self_attn.o_proj,0.0000024578,0.05000,28.644
|
| 246 |
+
34,self_attn.k_proj,0.0000380477,0.05000,28.668
|
| 247 |
+
35,mlp.gate_proj,0.0000984036,0.05000,18.177
|
| 248 |
+
35,mlp.up_proj,0.0000936725,0.05000,18.248
|
| 249 |
+
35,mlp.down_proj,0.0000177211,0.05000,21.932
|
| 250 |
+
35,self_attn.q_proj,0.0000328877,0.05000,27.925
|
| 251 |
+
35,self_attn.v_proj,0.0000387808,0.05000,28.576
|
| 252 |
+
35,self_attn.k_proj,0.0000347784,0.05000,28.628
|
| 253 |
+
35,self_attn.o_proj,0.0000015525,0.05000,28.713
|
| 254 |
+
36,mlp.gate_proj,0.0000940234,0.05000,18.255
|
| 255 |
+
36,mlp.up_proj,0.0000899822,0.05000,18.335
|
| 256 |
+
36,mlp.down_proj,0.0000180931,0.05000,21.839
|
| 257 |
+
36,self_attn.k_proj,0.0000334674,0.05000,29.467
|
| 258 |
+
36,self_attn.q_proj,0.0000326385,0.05000,29.651
|
| 259 |
+
36,self_attn.o_proj,0.0000020520,0.05000,29.665
|
| 260 |
+
36,self_attn.v_proj,0.0000368466,0.05000,29.685
|
| 261 |
+
37,mlp.gate_proj,0.0000945316,0.05000,19.052
|
| 262 |
+
37,mlp.up_proj,0.0000905586,0.05000,19.627
|
| 263 |
+
37,mlp.down_proj,0.0000239953,0.05000,23.262
|
| 264 |
+
37,self_attn.k_proj,0.0000336605,0.05000,29.512
|
| 265 |
+
37,self_attn.o_proj,0.0000015146,0.05000,29.655
|
| 266 |
+
37,self_attn.q_proj,0.0000327701,0.05000,29.723
|
| 267 |
+
37,self_attn.v_proj,0.0000377369,0.05000,29.729
|
| 268 |
+
38,mlp.up_proj,0.0000923737,0.05000,24.743
|
| 269 |
+
38,mlp.gate_proj,0.0000938228,0.05000,24.822
|
| 270 |
+
38,mlp.down_proj,0.0000328694,0.05000,28.743
|
| 271 |
+
38,self_attn.k_proj,0.0000334626,0.05000,29.978
|
| 272 |
+
38,self_attn.v_proj,0.0000377836,0.05000,30.020
|
| 273 |
+
38,self_attn.q_proj,0.0000331712,0.05000,30.023
|
| 274 |
+
38,self_attn.o_proj,0.0000035293,0.05000,30.060
|
| 275 |
+
39,mlp.gate_proj,0.0000792578,0.05000,24.641
|
| 276 |
+
39,mlp.up_proj,0.0000821907,0.05000,24.913
|
| 277 |
+
39,mlp.down_proj,0.0000582872,0.05000,28.858
|
| 278 |
+
39,self_attn.v_proj,0.0000294953,0.05000,29.453
|
| 279 |
+
39,self_attn.o_proj,0.0000085167,0.05000,29.570
|
| 280 |
+
39,self_attn.q_proj,0.0000305774,0.05000,29.662
|
| 281 |
+
39,self_attn.k_proj,0.0000303001,0.05000,29.671
|
quantize_config.json
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bits": 4,
|
| 3 |
+
"group_size": 128,
|
| 4 |
+
"desc_act": false,
|
| 5 |
+
"lm_head": false,
|
| 6 |
+
"method": "gptq",
|
| 7 |
+
"quant_method": "gptq",
|
| 8 |
+
"format": "gptq",
|
| 9 |
+
"checkpoint_format": "gptq",
|
| 10 |
+
"pack_dtype": "int32",
|
| 11 |
+
"meta": {
|
| 12 |
+
"quantizer": [
|
| 13 |
+
"gptqmodel:6.0.3"
|
| 14 |
+
],
|
| 15 |
+
"uri": "https://github.com/modelcloud/gptqmodel",
|
| 16 |
+
"damp_percent": 0.05,
|
| 17 |
+
"damp_auto_increment": 0.01,
|
| 18 |
+
"static_groups": false,
|
| 19 |
+
"true_sequential": true,
|
| 20 |
+
"mse": 0.0,
|
| 21 |
+
"gptaq": null,
|
| 22 |
+
"foem": null,
|
| 23 |
+
"act_group_aware": true,
|
| 24 |
+
"fallback": {
|
| 25 |
+
"strategy": "rtn",
|
| 26 |
+
"threshold": "0.5%",
|
| 27 |
+
"smooth": null
|
| 28 |
+
},
|
| 29 |
+
"offload_to_disk": true,
|
| 30 |
+
"offload_to_disk_path": "./gptqmodel_offload/ctrizrxh-tetrojst/",
|
| 31 |
+
"pack_impl": "cpu",
|
| 32 |
+
"gc_mode": "interval",
|
| 33 |
+
"wait_for_submodule_finalizers": false,
|
| 34 |
+
"auto_forward_data_parallel": true,
|
| 35 |
+
"vram_strategy": "exclusive",
|
| 36 |
+
"mock_quantization": false,
|
| 37 |
+
"hessian": {
|
| 38 |
+
"chunk_size": null,
|
| 39 |
+
"chunk_bytes": null,
|
| 40 |
+
"staging_dtype": "float32"
|
| 41 |
+
}
|
| 42 |
+
},
|
| 43 |
+
"sym": true
|
| 44 |
+
}
|
talkie_qmodel.py
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""GPTQModel adapter for the Talkie architecture.
|
| 2 |
+
|
| 3 |
+
Importing this module registers TalkieQModel under model_type='talkie' in
|
| 4 |
+
GPTQModel's MODEL_MAP, so `GPTQModel.load(...)` and `GPTQModel.from_quantized(...)`
|
| 5 |
+
work without manual configuration.
|
| 6 |
+
|
| 7 |
+
Auto-detect produces the same module_tree, so this is purely for the from_quantized
|
| 8 |
+
path (which doesn't run auto-detect — module_tree must be a class attribute).
|
| 9 |
+
"""
|
| 10 |
+
|
| 11 |
+
from __future__ import annotations
|
| 12 |
+
|
| 13 |
+
from gptqmodel.models.base import BaseQModel
|
| 14 |
+
from gptqmodel.models.auto import MODEL_MAP, SUPPORTED_MODELS
|
| 15 |
+
|
| 16 |
+
|
| 17 |
+
class TalkieQModel(BaseQModel):
|
| 18 |
+
# talkie uses functional F.rms_norm with no learnable scale, so there's no
|
| 19 |
+
# named pre-lm-head normalization module. Empty string disables that hook.
|
| 20 |
+
pre_lm_head_norm_module = ""
|
| 21 |
+
|
| 22 |
+
# Module tree maps GPTQModel's iteration onto our TalkieDecoderLayer:
|
| 23 |
+
# model.layers.{i}.self_attn.{q_proj,k_proj,v_proj,o_proj}
|
| 24 |
+
# model.layers.{i}.mlp.{gate_proj,up_proj,down_proj}
|
| 25 |
+
# Suffix :0/:1 declares quantization grouping order — q/k/v share input
|
| 26 |
+
# (the post-attn-rmsnorm hidden state), o has a different input (SDPA output).
|
| 27 |
+
# Same for gate/up sharing input vs down. Mirrors LlamaQModel.
|
| 28 |
+
module_tree = [
|
| 29 |
+
"model",
|
| 30 |
+
"layers",
|
| 31 |
+
"#",
|
| 32 |
+
{
|
| 33 |
+
"self_attn": ("q_proj:0", "k_proj:0", "v_proj:0", "o_proj:1"),
|
| 34 |
+
"mlp": ("gate_proj:0", "up_proj:0", "down_proj:1"),
|
| 35 |
+
},
|
| 36 |
+
]
|
| 37 |
+
|
| 38 |
+
|
| 39 |
+
# Register under model_type='talkie' so GPTQModel.load auto-routes to us.
|
| 40 |
+
MODEL_MAP["talkie"] = TalkieQModel
|
| 41 |
+
if "talkie" not in SUPPORTED_MODELS:
|
| 42 |
+
SUPPORTED_MODELS.append("talkie")
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:250b431ae085034dccfb9a4bf13db6bd1f8c375bb7ed499efd57e7c94ca64a3a
|
| 3 |
+
size 43187818
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"backend": "tokenizers",
|
| 3 |
+
"bos_token": null,
|
| 4 |
+
"clean_up_tokenization_spaces": false,
|
| 5 |
+
"eos_token": "<|endoftext|>",
|
| 6 |
+
"is_local": true,
|
| 7 |
+
"local_files_only": false,
|
| 8 |
+
"model_max_length": 1000000000000000019884624838656,
|
| 9 |
+
"pad_token": "<|endoftext|>",
|
| 10 |
+
"tokenizer_class": "TokenizersBackendFast",
|
| 11 |
+
"unk_token": null,
|
| 12 |
+
"_commit_hash": null
|
| 13 |
+
}
|