PEFT
Safetensors
Moshi
lora
personaplex
full-duplex
speech-to-speech
backchannel
maikezu commited on
Commit
2e6ab4d
·
verified ·
1 Parent(s): d9ead80

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +18 -148
README.md CHANGED
@@ -15,159 +15,29 @@ tags:
15
  - backchannel
16
  ---
17
 
18
- # personaplex-fisher-lora-backchannel-v14b
19
 
20
- ### Augmentation recipe
21
 
22
- | lever | value | effect |
23
- |---|---|---|
24
- | `--backchannel_aug_prob` | 0.3 | 30% of examples are augmented |
25
- | `--backchannel_aug_delete_share` | 1.0 | every augmented example is a *deletion* example — no insertions |
26
- | `--backchannel_aug_delete_prob` | 0.5 | half the backchannels in such an example are removed |
27
- | `--backchannel_aug_purge_prob` | 0.05 | 5% of examples have **every** backchannel erased, drawn independently of `aug_prob` |
28
 
29
- The purge is the important one. At `delete_prob 0.5` alone, every augmented example still contains
30
- backchannels, so the model can learn "fewer here" but never "none" — and "none" is the only signal
31
- that acts on *spontaneous* production. Purged examples keep their head positives, so head
32
- supervision is identical across all four v14 arms and the purge acts only on the LM.
33
 
34
- ## Training
35
 
36
- * **Data:** [`JSALT2026-Conv-AI-Simulator/fisher-v1`](https://huggingface.co/datasets/JSALT2026-Conv-AI-Simulator/fisher-v1)
37
- (Fisher telephone conversations), `max_length` 2000 frames (~160 s), 4 epochs, batch 1 x
38
- grad-accum 16.
39
- * **Backchannel labels** are derived from word alignments (Fisher's own
40
- `utterances[].backchannels` column is empty): each speaker is split into inter-pausal units, and
41
- one is marked a backchannel when it is short (`max_duration 1.5`), passes a content gate, and did
42
- not take the floor (`floor_gap 2.0`, retaining floor status across gaps up to 2 s).
43
- `--no-backchannel_include_laughter` — unlike `v2`, listener laughter does **not** count.
44
- `--backchannel_turn_clearance 1.5`, `--drop_non_speech_words`.
45
- * Only frames where a backchannel is *possible* (agent silent, partner talking,
46
- `agent_silence_margin 0.24`) are supervised; the rest are ignored rather than taught as negatives.
47
 
48
- ### Architecture
49
 
50
- * **LoRA:** attention only, on the temporal transformer —
51
- `^decoder\.model\.layers\.\d+\.(self_attn\.(q_proj|k_proj|v_proj|o_proj)\.linear)$` —
52
- `r=16`, `alpha=32`, `dropout=0.05`.
53
- * **Backchannel head:** `Linear(hidden -> 256) -> GELU -> Linear(256 -> 1)`, reading a single frame
54
- of the trunk's hidden state. `context=0`, i.e. **no** temporal convolution — that is the `claims`
55
- line of checkpoints.
56
- * **Trainable tokens:** the `<bc>` marker is token id **4**, the `<0x00>` byte-fallback slot — the
57
- one slot the tokenizer can never emit (it occurs zero times across all 11,396 Fisher rows). Its
58
- embedding row is trained in both `decoder.model.embed_tokens` and
59
- `depth_decoder.text_embed_tokens`; the rest of both tables stays frozen.
60
- * **Loss:** focal (`gamma=2.0`, `alpha=0.9`), backchannel loss weight `1.0`, head bias initialised
61
- to prior `0.016`, head LR `1e-3`, LM LR `2e-5`. Semantic codebook weight `1.0`, acoustic `0.02`,
62
- pad-text loss weight `0.3`.
63
-
64
- ## Evaluation
65
-
66
- Three measurements that answer three different questions:
67
-
68
- | measurement | what it asks | material |
69
- |---|---|---|
70
- | Fisher head eval | can the head rank frames? | Fisher validation, teacher-forced |
71
- | TurnBench head eval | does that survive out of domain, against human annotations? | TurnBench dev, teacher-forced |
72
- | self-play obedience | when the head fires, does the model produce a backchannel? | Fisher-domain self-play |
73
-
74
- Matching tolerance is 5 frames (400 ms) in both head evals — exact-frame agreement is too strict a
75
- bar for a phenomenon whose natural placement window is a few hundred ms wide.
76
-
77
- ### Stage 1 — Fisher validation, teacher-forced
78
-
79
- 93,042 supervised frames, 1,213 positive (1.30%).
80
-
81
- | Average Precision | ROC-AUC | mean p on positives | mean p on negatives |
82
- |---|---|---|---|
83
- | 0.352 | 0.928 | 0.486 | 0.182 |
84
-
85
- | Threshold | Precision | Recall | F1@400ms | Rate vs. human |
86
- |---|---|---|---|---|
87
- | 0.70 | 0.206 | 0.305 | 0.246 | 1.48x |
88
-
89
-
90
- ### Stage 2 — TurnBench dev, out of domain
91
-
92
- 12 conversations / 24 dialogue passes, 6,553 supervised frames, 97 positive, 35 reference events,
93
- human rate 1.52 events/min. Run on a **-23 LUFS loudness-normalised** copy: TurnBench ships 19.4 dB
94
- apart internally and ~15 dB below Fisher, which moves where the head fires.
95
-
96
- Two conditions: `keep` leaves the recording untouched (the agent's own backchannels stay in its
97
- context, so the head can lean on them); `remove` erases them, which is the condition that matches
98
- self-play and the render harness.
99
-
100
- | Condition | AP | ROC-AUC | Best lag | Rate-matched thr | P | R | **F1** | Rate |
101
- |---|---|---|---|---|---|---|---|---|
102
- | `keep` | 0.147 | 0.810 | -1 | 0.650 | 0.222 | 0.229 | **0.225** | 1.03x |
103
- | `remove` | 0.123 | 0.805 | -4 | 0.650 | 0.485 | 0.457 | **0.471** | 0.94x |
104
-
105
- At best-F1 rather than rate-matched, `remove` reaches **F1 0.475** (thr 0.675, P 0.583 / R 0.400) but
106
- sits at 0.69x the human rate — it buys precision by firing less.
107
-
108
- Two caveats worth carrying:
109
-
110
- * **The in-domain to out-of-domain drop is large and not checkpoint-specific**: AP 0.35 -> 0.12-0.15,
111
- ROC-AUC 0.93 -> 0.81.
112
- * **v14b's optimal lag on `remove` is -4 frames** (320 ms *before* the human onset), where the other
113
- arms sit at -1/-2. AP at that lag is 0.231, against 0.123 at lag 0.
114
-
115
- ### Stage 3 — self-play obedience
116
-
117
- 100 dialogues x 80 s windows (133 min) in the partner-backchannelling condition. Thresholds are
118
- calibrated **on generated streams**, never reused from validation — confidence is systematically
119
- different in free generation than under teacher forcing.
120
-
121
- | Condition | Threshold | Fires/min | **Obeyed** | **Realised** | clean bc | overrun | other | stranded |
122
- |---|---|---|---|---|---|---|---|---|
123
- | off | — | 0.00 | — | — | — | — | — | — |
124
- | low | 0.740 | 1.03 | **100%** | **77%** | 106 | 31 | 0 | 0 |
125
- | normal | 0.695 | 1.48 | **98%** | **69%** | 137 | 58 | 3 | 0 |
126
- | high | 0.615 | 3.59 | **95%** | **65%** | 311 | 142 | 24 | 2 |
127
- | max | 0.485 | 8.31 | **89%** | **65%** | 717 | 270 | 84 | 37 |
128
- | *Partner (within-dialogue baseline)* | — | *1.51* | — | — | — | — | — | — |
129
-
130
- * **`obeyed`** = the model started *something* at the trigger. **`realised`** = what it said was a
131
- backchannel and stayed one. **`overrun`** = it began with a backchannel and then kept talking into
132
- a turn. **`other`** = not a backchannel at all.
133
- * At `normal`, only **3 of 198** triggers produced a non-backchannel.
134
- * Obedience degrades gracefully: still 89% obeyed when driven to 5.5x the human rate.
135
- * **Realised forms** are the expected inventory: `yeah` (~500), `right`, `okay`, `uh-huh`, `mm`,
136
- `mhm`.
137
-
138
- #### Controllability of `off`
139
-
140
- With the head off, **zero** forced events fire, but the model still emits **0.70 spontaneous
141
- lexical backchannels/min**, against 1.51/min for its partner in the same calls. So `off` means "no
142
- head-driven backchannels", not silence — the purge lever moves this number but does not zero it.
143
-
144
- #### Overrun
145
-
146
- 29% of obeyed triggers at `normal` (58/198) overran into a turn. The forced-decode harness measures
147
- the same phenomenon far lower — 22/177 = 12% continued past a complete backchannel, of which 9
148
- (5.1%) were true turn attempts — because its gate requires 3.2 s of clearance before firing, so it
149
- only fires where a turn *cannot* start. Self-play lacks that constraint, so **treat 29% as the
150
- pessimistic bound and 5% as what a clearance-gated deployment sees.**
151
-
152
- ### Choosing among the v14 arms
153
-
154
- Rate-matched F1 on TurnBench, and obedience at a matched ~1.5 fires/min:
155
-
156
- | ckpt | augmentation | Fisher F1 | TB `keep` F1 | TB `remove` F1 | obeyed | realised |
157
- |---|---|---|---|---|---|---|
158
- | v14a | none | .252 | **.267** | .348 | 87% | 51% |
159
- | **v14b** | **deletions only** | .246 | .225 | .471 | **98%** | **69%** |
160
- | v14c | insertions only | .244 | **.270** | .324 | 93% | 63% |
161
- | v14d | both | .246 | .257 | **.479** | 93% | 53% |
162
-
163
- 1. **Fisher separates nothing** — all four within 0.008 F1 and 0.002 AUC on 1,213 positive frames.
164
- 2. **`remove` splits them into two pairs**: v14b/v14d at .471/.479 against v14a/v14c at .348/.324, a
165
- ~45% relative gap driven by precision. But v14b vs v14d is a coin flip on 35 reference events.
166
- 3. **`keep` ranks them differently** and puts v14b last. `keep` is the easier condition and the one
167
- that does not match how the model is actually run.
168
- 4. **Obedience is where they separate, and v14b wins outright** — at every floor, on both metrics,
169
- with 3 non-backchannels out of 198 against 26 for v14a.
170
-
171
- Counterintuitive and worth restating: **deletions-only beats insertions-only and both**, even though
172
- insertion is the arm meant to teach the model to obey a backchannel request.
173
 
 
 
 
 
 
 
 
 
 
 
 
 
15
  - backchannel
16
  ---
17
 
 
18
 
19
+ # PersonaPlex Backchannel Head
20
 
21
+ This HF repository contains a LoRA adapter for [PersonaPlex-7B](https://huggingface.co/nvidia/personaplex-7b-v1) that adds a lightweight, controllable backchannel head, introduced in [Controlling Backchannels in Streamable Full-Duplex Models](https://arxiv.org/abs/2609.29418).
 
 
 
 
 
22
 
23
+ About our work: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.
 
 
 
24
 
25
+ Please refer to the [codebase](https://github.com/MaikeZuefle/bcmore) for the usage of the model.
26
 
27
+ For more information, please have a look at the paper.
 
 
 
 
 
 
 
 
 
 
28
 
29
+ ## Citation
30
 
31
+ If you use this model, please cite:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
 
33
+ ```bibtex
34
+ @misc{züfle2026controllingbackchannelsstreamablefullduplex,
35
+ title={Controlling Backchannels in Streamable Full-duplex Models},
36
+ author={Maike Züfle and Peter Polák and Sefik Emre Eskimez and Jan Niehues and Peter Bell and Ondřej Klejch},
37
+ year={2026},
38
+ eprint={2609.29418},
39
+ archivePrefix={arXiv},
40
+ primaryClass={cs.CL},
41
+ url={https://arxiv.org/abs/2609.29418},
42
+ }
43
+ ```