adidsh commited on
Commit
6556a52
Β·
verified Β·
1 Parent(s): 8abdbe2

Upload token_contract.md

Browse files
Files changed (1) hide show
  1. token_contract.md +197 -0
token_contract.md ADDED
@@ -0,0 +1,197 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Token contract β€” `llama-3-audio-tokenizer`
2
+
3
+ The authoritative description of what every token ID **means** in this project's
4
+ compiled data, checkpoints and eval paths. Generated from the tokenizer itself,
5
+ not transcribed by hand.
6
+
7
+ tokenizer: checkpoints/llama-3-audio-tokenizer
8
+ vocab size: 156,960
9
+ fingerprint: b645cf6612315393eea21d29e94304bdbcbfee131a25de5eecf6bae6ee9f39c5
10
+ base model: Llama-3.2-3B (128,000 BPE + 256 Llama specials)
11
+
12
+ **A token ID is meaningless without this contract.** Compiled parquet stores raw
13
+ integers; nothing in the data says which tokenizer minted them. Compiling with one
14
+ tokenizer and training with another is silent β€” training runs, loss descends, and
15
+ the model learns the wrong symbol for every audio frame. That is not hypothetical:
16
+ it is exactly how the gemma3 SNAC base mismatch produced undecodable checkpoints,
17
+ caught only at eval. Hence the fingerprint, stamped at compile time and asserted at
18
+ training startup and in the eval decode path (`scripts/token_contract.py`).
19
+
20
+ ---
21
+
22
+ ## 1. ID map
23
+
24
+ | range | count | what it is |
25
+ |---|---:|---|
26
+ | `0 – 127,999` | 128,000 | Llama-3 BPE text vocabulary |
27
+ | `128,000 – 128,255` | 256 | stock Llama-3 specials (`<\|begin_of_text\|>`, `<\|eot_id\|>`, `<\|reserved_special_token_N\|>` …) |
28
+ | `128,256 – 128,265` | 10 | **project control tokens** (Β§2) |
29
+ | `128,266 – 156,937` | 28,672 | **SNAC audio codes** β€” 7 codebooks Γ— 4,096 (Β§3) |
30
+ | `156,938 – 156,959` | 22 | **conditioning + paralinguistic tokens** (Β§4) |
31
+
32
+ 28,960 added tokens in total (28,672 SNAC + 288 non-SNAC).
33
+
34
+ Note the layout is *discontinuous*: the control block sits immediately below the
35
+ audio band and the conditioning block immediately above it. Anything that assumes
36
+ "all added tokens are contiguous above the base" is wrong.
37
+
38
+ ---
39
+
40
+ ## 2. Control tokens β€” `128,256 – 128,265`
41
+
42
+ | id | token | role |
43
+ |---:|---|---|
44
+ | 128256 | `<\|reserved_0\|>` | unused |
45
+ | **128257** | `<\|start_of_speech\|>` | opens the audio span; the model emits this itself as its first generated token |
46
+ | **128258** | `<\|end_of_speech\|>` | closes the audio span |
47
+ | **128259** | `<\|start_of_human\|>` | opens the prompt turn |
48
+ | **128260** | `<\|end_of_human\|>` | closes the prompt turn |
49
+ | **128261** | `<\|start_of_ai\|>` | opens the model turn |
50
+ | **128262** | `<\|end_of_ai\|>` | closes the model turn |
51
+ | 128263 | `<\|pad\|>` | padding |
52
+ | 128264 | `<\|reserved_8\|>` | unused |
53
+ | 128265 | `<\|reserved_9\|>` | unused |
54
+
55
+ ---
56
+
57
+ ## 3. Audio band β€” `128,266 – 156,937`
58
+
59
+ audio_token_base_id = 128266
60
+ codebooks = 7
61
+ codebook size = 4096
62
+ total audio tokens = 28672 (base .. base + 7*4096 - 1 = 156937)
63
+
64
+ SNAC 24 kHz emits three codebooks at a 1:2:4 temporal ratio:
65
+
66
+ c0: [seq_len] 12 Hz coarsest
67
+ c1: [2*seq_len] 23 Hz
68
+ c2: [4*seq_len] 47 Hz finest
69
+
70
+ These are flattened to **7 tokens per frame**, in this fixed interleave:
71
+
72
+ frame i -> [ c0[i], c1[2i], c2[4i], c2[4i+1], c1[2i+1], c2[4i+2], c2[4i+3] ]
73
+
74
+ Each raw code (0–4095) is offset by its **position in the frame**, not by which
75
+ codebook it came from:
76
+
77
+ | frame position | source | offset | id range |
78
+ |---:|---|---|---|
79
+ | 0 | `c0[i]` | base + 0Γ—4096 | 128,266 – 132,361 |
80
+ | 1 | `c1[2i]` | base + 1Γ—4096 | 132,362 – 136,457 |
81
+ | 2 | `c2[4i]` | base + 2Γ—4096 | 136,458 – 140,553 |
82
+ | 3 | `c2[4i+1]` | base + 3Γ—4096 | 140,554 – 144,649 |
83
+ | 4 | `c1[2i+1]` | base + 4Γ—4096 | 144,650 – 148,745 |
84
+ | 5 | `c2[4i+2]` | base + 5Γ—4096 | 148,746 – 152,841 |
85
+ | 6 | `c2[4i+3]` | base + 6Γ—4096 | 152,842 – 156,937 |
86
+
87
+ So positions 1 and 4 are both `c1`, and 2/3/5/6 are all `c2` β€” the offset encodes
88
+ *where in the frame* a code sits, which is what makes the stream decodable without
89
+ a separate structure signal.
90
+
91
+ **Consecutive duplicate frames are removed** at encode time (frames sharing the
92
+ same `c0`), so token count is not exactly proportional to duration.
93
+
94
+ Derived constants (`scripts/snac_tokenizer.py`):
95
+
96
+ samples per c0 frame = 512 (encoder_rates [2,4,8,8] x vq_stride 4)
97
+ streaming window = 4 frames = 28 tokens, middle frame kept
98
+ long-audio window = 512 frames (~43.7 s), 4-frame context each side
99
+
100
+ Rule of thumb used throughout the project: **82 tokens β‰ˆ 1 second of audio.**
101
+
102
+ ---
103
+
104
+ ## 4. Conditioning + paralinguistics β€” `156,938 – 156,959`
105
+
106
+ | id | token | role |
107
+ |---:|---|---|
108
+ | 156938 / 156939 | `<\|speaker>` / `<speaker\|>` | wrap a speaker name |
109
+ | 156940 / 156941 | `<\|style>` / `<style\|>` | wrap a style label (rasa) or accent string (globe) |
110
+ | 156942 / 156943 | `<\|env>` / `<env\|>` | wrap an environment label; open class, label stays free text |
111
+
112
+ Non-verbals β€” a **closed set of 16**, one token each (no wrapper):
113
+
114
+ | id | token | | id | token |
115
+ |---:|---|---|---:|---|
116
+ | 156944 | `<\|nv_breath\|>` | | 156952 | `<\|nv_hum\|>` |
117
+ | 156945 | `<\|nv_stammer\|>` | | 156953 | `<\|nv_gasp\|>` |
118
+ | 156946 | `<\|nv_throat\|>` | | 156954 | `<\|nv_wheeze\|>` |
119
+ | 156947 | `<\|nv_laugh\|>` | | 156955 | `<\|nv_sneeze\|>` |
120
+ | 156948 | `<\|nv_swallow\|>` | | 156956 | `<\|nv_snort\|>` |
121
+ | 156949 | `<\|nv_sniff\|>` | | 156957 | `<\|nv_yawn\|>` |
122
+ | 156950 | `<\|nv_sigh\|>` | | 156958 | `<\|nv_groan\|>` |
123
+ | 156951 | `<\|nv_cough\|>` | | 156959 | `<\|nv_burp\|>` |
124
+
125
+ The asymmetry is deliberate: non-verbals are a closed vocabulary and get dedicated
126
+ tokens; environment labels are an open class, so only the wrapper is a token and
127
+ the label inside stays free text β€” a new environment class needs no tokenizer
128
+ change. Source-form mapping (`<breath>`, `[bird_squawk]`, …) is in
129
+ `scripts/paralinguistics.py`.
130
+
131
+ ---
132
+
133
+ ## 5. Sequence layouts
134
+
135
+ Built by `scripts/chat_templates.py`. `prompt_end` is defined as everything up to
136
+ and **including** `<|start_of_ai|>` β€” the prompt therefore stops *before*
137
+ `<|start_of_speech|>`, which the model emits itself.
138
+
139
+ **TTS with conditioning** (rasa, globe, bhili):
140
+
141
+ <|start_of_human|><|begin_of_text|>
142
+ <|speaker>NAME<speaker|>\n
143
+ <|style>LABEL<style|>\n
144
+ TEXT
145
+ <|eot_id|><|end_of_human|><|start_of_ai|>
146
+ <|start_of_speech|> ...audio... <|end_of_speech|>
147
+ <|end_of_ai|>
148
+
149
+ Metadata order is always **speaker β†’ style β†’ accent**, each block followed by a
150
+ newline. An empty value emits nothing at all (no empty wrapper).
151
+
152
+ **Multi-turn conversation** (`gemini_vc_conversational`) β€” no metadata prefix;
153
+ speaker labels are inline turn markers inside the text:
154
+
155
+ <|start_of_human|><|begin_of_text|>
156
+ <|speaker>A<speaker|>\nturn one\n\n<|speaker>B<speaker|>\nturn two ...
157
+ <|eot_id|><|end_of_human|><|start_of_ai|><|start_of_speech|> ...
158
+
159
+ Turn separator is a **blank line** (`\n\n`) before each subsequent
160
+ `<|speaker>` marker. `normalize_text` preserves newline runs (the old
161
+ collapse-to-one-space behavior was removed when the `\n\n` issue was fixed),
162
+ so the separator reaches the model verbatim; verified present in 100% of
163
+ compiled multi-turn rows in both `gemini_vc` and `gemini_src_conv`. The `\n`
164
+ after `<speaker|>` also survives. Inference prompts must match this form.
165
+
166
+ ---
167
+
168
+ ## 6. Fingerprint
169
+
170
+ `scripts/token_contract.py::compute_fingerprint` is a sha256 over exactly the
171
+ things that change what an ID means:
172
+
173
+ - the full added-token map (content β†’ id), sorted
174
+ - vocab size
175
+ - the SNAC base id and layout constants (codebooks Γ— codebook size)
176
+
177
+ It deliberately **excludes** `tokenizer_config.json` niceties β€” padding side, chat
178
+ template, `model_max_length` β€” because those do not change a token's meaning and
179
+ including them would fire on cosmetic edits.
180
+
181
+ current fingerprint: b645cf6612315393eea21d29e94304bdbcbfee131a25de5eecf6bae6ee9f39c5
182
+
183
+ Why a content hash and not a range check: the audio-band range check in
184
+ `snac_tokenizer.decode_audio` catches the loud failure (IDs outside the band). It
185
+ cannot catch the quiet one β€” sibling tokenizers `llama-3-audio-tokenizer`
186
+ (156,938), `-tok_trimmed` (156,942) and `-style` (156,952) all share base 128,266,
187
+ so every range check passes while `<|style>` means something different in the data
188
+ than in the model.
189
+
190
+ Mismatch is a hard error, never a warning. A silent wrong-tokenizer run costs a
191
+ full training cycle.
192
+
193
+ ---
194
+
195
+ *Generated from `checkpoints/llama-3-audio-tokenizer` with `scripts/token_contract.py`;
196
+ layout constants from `scripts/snac_tokenizer.py`, sequence templates from
197
+ `scripts/chat_templates.py`.*