PEFT
Safetensors
Moshi
lora
personaplex
full-duplex
speech-to-speech
backchannel
peterpo commited on
Commit
d9ead80
Β·
verified Β·
1 Parent(s): e1266b7

Add model card

Browse files
Files changed (1) hide show
  1. README.md +173 -0
README.md ADDED
@@ -0,0 +1,173 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: nvidia/personaplex-7b-v1
3
+ library_name: peft
4
+ license: other
5
+ license_name: nvidia-open-model-license
6
+ license_link: https://huggingface.co/nvidia/personaplex-7b-v1
7
+ datasets:
8
+ - JSALT2026-Conv-AI-Simulator/fisher-v1
9
+ tags:
10
+ - lora
11
+ - personaplex
12
+ - moshi
13
+ - full-duplex
14
+ - speech-to-speech
15
+ - backchannel
16
+ ---
17
+
18
+ # personaplex-fisher-lora-backchannel-v14b
19
+
20
+ ### Augmentation recipe
21
+
22
+ | lever | value | effect |
23
+ |---|---|---|
24
+ | `--backchannel_aug_prob` | 0.3 | 30% of examples are augmented |
25
+ | `--backchannel_aug_delete_share` | 1.0 | every augmented example is a *deletion* example β€” no insertions |
26
+ | `--backchannel_aug_delete_prob` | 0.5 | half the backchannels in such an example are removed |
27
+ | `--backchannel_aug_purge_prob` | 0.05 | 5% of examples have **every** backchannel erased, drawn independently of `aug_prob` |
28
+
29
+ The purge is the important one. At `delete_prob 0.5` alone, every augmented example still contains
30
+ backchannels, so the model can learn "fewer here" but never "none" β€” and "none" is the only signal
31
+ that acts on *spontaneous* production. Purged examples keep their head positives, so head
32
+ supervision is identical across all four v14 arms and the purge acts only on the LM.
33
+
34
+ ## Training
35
+
36
+ * **Data:** [`JSALT2026-Conv-AI-Simulator/fisher-v1`](https://huggingface.co/datasets/JSALT2026-Conv-AI-Simulator/fisher-v1)
37
+ (Fisher telephone conversations), `max_length` 2000 frames (~160 s), 4 epochs, batch 1 x
38
+ grad-accum 16.
39
+ * **Backchannel labels** are derived from word alignments (Fisher's own
40
+ `utterances[].backchannels` column is empty): each speaker is split into inter-pausal units, and
41
+ one is marked a backchannel when it is short (`max_duration 1.5`), passes a content gate, and did
42
+ not take the floor (`floor_gap 2.0`, retaining floor status across gaps up to 2 s).
43
+ `--no-backchannel_include_laughter` β€” unlike `v2`, listener laughter does **not** count.
44
+ `--backchannel_turn_clearance 1.5`, `--drop_non_speech_words`.
45
+ * Only frames where a backchannel is *possible* (agent silent, partner talking,
46
+ `agent_silence_margin 0.24`) are supervised; the rest are ignored rather than taught as negatives.
47
+
48
+ ### Architecture
49
+
50
+ * **LoRA:** attention only, on the temporal transformer β€”
51
+ `^decoder\.model\.layers\.\d+\.(self_attn\.(q_proj|k_proj|v_proj|o_proj)\.linear)$` β€”
52
+ `r=16`, `alpha=32`, `dropout=0.05`.
53
+ * **Backchannel head:** `Linear(hidden -> 256) -> GELU -> Linear(256 -> 1)`, reading a single frame
54
+ of the trunk's hidden state. `context=0`, i.e. **no** temporal convolution β€” that is the `claims`
55
+ line of checkpoints.
56
+ * **Trainable tokens:** the `<bc>` marker is token id **4**, the `<0x00>` byte-fallback slot β€” the
57
+ one slot the tokenizer can never emit (it occurs zero times across all 11,396 Fisher rows). Its
58
+ embedding row is trained in both `decoder.model.embed_tokens` and
59
+ `depth_decoder.text_embed_tokens`; the rest of both tables stays frozen.
60
+ * **Loss:** focal (`gamma=2.0`, `alpha=0.9`), backchannel loss weight `1.0`, head bias initialised
61
+ to prior `0.016`, head LR `1e-3`, LM LR `2e-5`. Semantic codebook weight `1.0`, acoustic `0.02`,
62
+ pad-text loss weight `0.3`.
63
+
64
+ ## Evaluation
65
+
66
+ Three measurements that answer three different questions:
67
+
68
+ | measurement | what it asks | material |
69
+ |---|---|---|
70
+ | Fisher head eval | can the head rank frames? | Fisher validation, teacher-forced |
71
+ | TurnBench head eval | does that survive out of domain, against human annotations? | TurnBench dev, teacher-forced |
72
+ | self-play obedience | when the head fires, does the model produce a backchannel? | Fisher-domain self-play |
73
+
74
+ Matching tolerance is 5 frames (400 ms) in both head evals β€” exact-frame agreement is too strict a
75
+ bar for a phenomenon whose natural placement window is a few hundred ms wide.
76
+
77
+ ### Stage 1 β€” Fisher validation, teacher-forced
78
+
79
+ 93,042 supervised frames, 1,213 positive (1.30%).
80
+
81
+ | Average Precision | ROC-AUC | mean p on positives | mean p on negatives |
82
+ |---|---|---|---|
83
+ | 0.352 | 0.928 | 0.486 | 0.182 |
84
+
85
+ | Threshold | Precision | Recall | F1@400ms | Rate vs. human |
86
+ |---|---|---|---|---|
87
+ | 0.70 | 0.206 | 0.305 | 0.246 | 1.48x |
88
+
89
+
90
+ ### Stage 2 β€” TurnBench dev, out of domain
91
+
92
+ 12 conversations / 24 dialogue passes, 6,553 supervised frames, 97 positive, 35 reference events,
93
+ human rate 1.52 events/min. Run on a **-23 LUFS loudness-normalised** copy: TurnBench ships 19.4 dB
94
+ apart internally and ~15 dB below Fisher, which moves where the head fires.
95
+
96
+ Two conditions: `keep` leaves the recording untouched (the agent's own backchannels stay in its
97
+ context, so the head can lean on them); `remove` erases them, which is the condition that matches
98
+ self-play and the render harness.
99
+
100
+ | Condition | AP | ROC-AUC | Best lag | Rate-matched thr | P | R | **F1** | Rate |
101
+ |---|---|---|---|---|---|---|---|---|
102
+ | `keep` | 0.147 | 0.810 | -1 | 0.650 | 0.222 | 0.229 | **0.225** | 1.03x |
103
+ | `remove` | 0.123 | 0.805 | -4 | 0.650 | 0.485 | 0.457 | **0.471** | 0.94x |
104
+
105
+ At best-F1 rather than rate-matched, `remove` reaches **F1 0.475** (thr 0.675, P 0.583 / R 0.400) but
106
+ sits at 0.69x the human rate β€” it buys precision by firing less.
107
+
108
+ Two caveats worth carrying:
109
+
110
+ * **The in-domain to out-of-domain drop is large and not checkpoint-specific**: AP 0.35 -> 0.12-0.15,
111
+ ROC-AUC 0.93 -> 0.81.
112
+ * **v14b's optimal lag on `remove` is -4 frames** (320 ms *before* the human onset), where the other
113
+ arms sit at -1/-2. AP at that lag is 0.231, against 0.123 at lag 0.
114
+
115
+ ### Stage 3 β€” self-play obedience
116
+
117
+ 100 dialogues x 80 s windows (133 min) in the partner-backchannelling condition. Thresholds are
118
+ calibrated **on generated streams**, never reused from validation β€” confidence is systematically
119
+ different in free generation than under teacher forcing.
120
+
121
+ | Condition | Threshold | Fires/min | **Obeyed** | **Realised** | clean bc | overrun | other | stranded |
122
+ |---|---|---|---|---|---|---|---|---|
123
+ | off | β€” | 0.00 | β€” | β€” | β€” | β€” | β€” | β€” |
124
+ | low | 0.740 | 1.03 | **100%** | **77%** | 106 | 31 | 0 | 0 |
125
+ | normal | 0.695 | 1.48 | **98%** | **69%** | 137 | 58 | 3 | 0 |
126
+ | high | 0.615 | 3.59 | **95%** | **65%** | 311 | 142 | 24 | 2 |
127
+ | max | 0.485 | 8.31 | **89%** | **65%** | 717 | 270 | 84 | 37 |
128
+ | *Partner (within-dialogue baseline)* | β€” | *1.51* | β€” | β€” | β€” | β€” | β€” | β€” |
129
+
130
+ * **`obeyed`** = the model started *something* at the trigger. **`realised`** = what it said was a
131
+ backchannel and stayed one. **`overrun`** = it began with a backchannel and then kept talking into
132
+ a turn. **`other`** = not a backchannel at all.
133
+ * At `normal`, only **3 of 198** triggers produced a non-backchannel.
134
+ * Obedience degrades gracefully: still 89% obeyed when driven to 5.5x the human rate.
135
+ * **Realised forms** are the expected inventory: `yeah` (~500), `right`, `okay`, `uh-huh`, `mm`,
136
+ `mhm`.
137
+
138
+ #### Controllability of `off`
139
+
140
+ With the head off, **zero** forced events fire, but the model still emits **0.70 spontaneous
141
+ lexical backchannels/min**, against 1.51/min for its partner in the same calls. So `off` means "no
142
+ head-driven backchannels", not silence β€” the purge lever moves this number but does not zero it.
143
+
144
+ #### Overrun
145
+
146
+ 29% of obeyed triggers at `normal` (58/198) overran into a turn. The forced-decode harness measures
147
+ the same phenomenon far lower β€” 22/177 = 12% continued past a complete backchannel, of which 9
148
+ (5.1%) were true turn attempts β€” because its gate requires 3.2 s of clearance before firing, so it
149
+ only fires where a turn *cannot* start. Self-play lacks that constraint, so **treat 29% as the
150
+ pessimistic bound and 5% as what a clearance-gated deployment sees.**
151
+
152
+ ### Choosing among the v14 arms
153
+
154
+ Rate-matched F1 on TurnBench, and obedience at a matched ~1.5 fires/min:
155
+
156
+ | ckpt | augmentation | Fisher F1 | TB `keep` F1 | TB `remove` F1 | obeyed | realised |
157
+ |---|---|---|---|---|---|---|
158
+ | v14a | none | .252 | **.267** | .348 | 87% | 51% |
159
+ | **v14b** | **deletions only** | .246 | .225 | .471 | **98%** | **69%** |
160
+ | v14c | insertions only | .244 | **.270** | .324 | 93% | 63% |
161
+ | v14d | both | .246 | .257 | **.479** | 93% | 53% |
162
+
163
+ 1. **Fisher separates nothing** β€” all four within 0.008 F1 and 0.002 AUC on 1,213 positive frames.
164
+ 2. **`remove` splits them into two pairs**: v14b/v14d at .471/.479 against v14a/v14c at .348/.324, a
165
+ ~45% relative gap driven by precision. But v14b vs v14d is a coin flip on 35 reference events.
166
+ 3. **`keep` ranks them differently** and puts v14b last. `keep` is the easier condition and the one
167
+ that does not match how the model is actually run.
168
+ 4. **Obedience is where they separate, and v14b wins outright** β€” at every floor, on both metrics,
169
+ with 3 non-backchannels out of 198 against 26 for v14a.
170
+
171
+ Counterintuitive and worth restating: **deletions-only beats insertions-only and both**, even though
172
+ insertion is the arm meant to teach the model to obey a backchannel request.
173
+