Instructions to use JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head with PEFT:
Task type is invalid.
- Moshi
How to use JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head with Moshi:
# pip install moshi # Run the interactive web server python -m moshi.server --hf-repo "JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head" # Then open https://localhost:8998 in your browser
# pip install moshi import torch from moshi.models import loaders # Load checkpoint info from HuggingFace checkpoint = loaders.CheckpointInfo.from_hf_repo("JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head") # Load the Mimi audio codec mimi = checkpoint.get_mimi(device="cuda") mimi.set_num_codebooks(8) # Encode audio (24kHz, mono) wav = torch.randn(1, 1, 24000 * 10) # [batch, channels, samples] with torch.no_grad(): codes = mimi.encode(wav.cuda()) decoded = mimi.decode(codes) - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -15,159 +15,29 @@ tags:
|
|
| 15 |
- backchannel
|
| 16 |
---
|
| 17 |
|
| 18 |
-
# personaplex-fisher-lora-backchannel-v14b
|
| 19 |
|
| 20 |
-
#
|
| 21 |
|
| 22 |
-
|
| 23 |
-
|---|---|---|
|
| 24 |
-
| `--backchannel_aug_prob` | 0.3 | 30% of examples are augmented |
|
| 25 |
-
| `--backchannel_aug_delete_share` | 1.0 | every augmented example is a *deletion* example — no insertions |
|
| 26 |
-
| `--backchannel_aug_delete_prob` | 0.5 | half the backchannels in such an example are removed |
|
| 27 |
-
| `--backchannel_aug_purge_prob` | 0.05 | 5% of examples have **every** backchannel erased, drawn independently of `aug_prob` |
|
| 28 |
|
| 29 |
-
|
| 30 |
-
backchannels, so the model can learn "fewer here" but never "none" — and "none" is the only signal
|
| 31 |
-
that acts on *spontaneous* production. Purged examples keep their head positives, so head
|
| 32 |
-
supervision is identical across all four v14 arms and the purge acts only on the LM.
|
| 33 |
|
| 34 |
-
|
| 35 |
|
| 36 |
-
|
| 37 |
-
(Fisher telephone conversations), `max_length` 2000 frames (~160 s), 4 epochs, batch 1 x
|
| 38 |
-
grad-accum 16.
|
| 39 |
-
* **Backchannel labels** are derived from word alignments (Fisher's own
|
| 40 |
-
`utterances[].backchannels` column is empty): each speaker is split into inter-pausal units, and
|
| 41 |
-
one is marked a backchannel when it is short (`max_duration 1.5`), passes a content gate, and did
|
| 42 |
-
not take the floor (`floor_gap 2.0`, retaining floor status across gaps up to 2 s).
|
| 43 |
-
`--no-backchannel_include_laughter` — unlike `v2`, listener laughter does **not** count.
|
| 44 |
-
`--backchannel_turn_clearance 1.5`, `--drop_non_speech_words`.
|
| 45 |
-
* Only frames where a backchannel is *possible* (agent silent, partner talking,
|
| 46 |
-
`agent_silence_margin 0.24`) are supervised; the rest are ignored rather than taught as negatives.
|
| 47 |
|
| 48 |
-
##
|
| 49 |
|
| 50 |
-
|
| 51 |
-
`^decoder\.model\.layers\.\d+\.(self_attn\.(q_proj|k_proj|v_proj|o_proj)\.linear)$` —
|
| 52 |
-
`r=16`, `alpha=32`, `dropout=0.05`.
|
| 53 |
-
* **Backchannel head:** `Linear(hidden -> 256) -> GELU -> Linear(256 -> 1)`, reading a single frame
|
| 54 |
-
of the trunk's hidden state. `context=0`, i.e. **no** temporal convolution — that is the `claims`
|
| 55 |
-
line of checkpoints.
|
| 56 |
-
* **Trainable tokens:** the `<bc>` marker is token id **4**, the `<0x00>` byte-fallback slot — the
|
| 57 |
-
one slot the tokenizer can never emit (it occurs zero times across all 11,396 Fisher rows). Its
|
| 58 |
-
embedding row is trained in both `decoder.model.embed_tokens` and
|
| 59 |
-
`depth_decoder.text_embed_tokens`; the rest of both tables stays frozen.
|
| 60 |
-
* **Loss:** focal (`gamma=2.0`, `alpha=0.9`), backchannel loss weight `1.0`, head bias initialised
|
| 61 |
-
to prior `0.016`, head LR `1e-3`, LM LR `2e-5`. Semantic codebook weight `1.0`, acoustic `0.02`,
|
| 62 |
-
pad-text loss weight `0.3`.
|
| 63 |
-
|
| 64 |
-
## Evaluation
|
| 65 |
-
|
| 66 |
-
Three measurements that answer three different questions:
|
| 67 |
-
|
| 68 |
-
| measurement | what it asks | material |
|
| 69 |
-
|---|---|---|
|
| 70 |
-
| Fisher head eval | can the head rank frames? | Fisher validation, teacher-forced |
|
| 71 |
-
| TurnBench head eval | does that survive out of domain, against human annotations? | TurnBench dev, teacher-forced |
|
| 72 |
-
| self-play obedience | when the head fires, does the model produce a backchannel? | Fisher-domain self-play |
|
| 73 |
-
|
| 74 |
-
Matching tolerance is 5 frames (400 ms) in both head evals — exact-frame agreement is too strict a
|
| 75 |
-
bar for a phenomenon whose natural placement window is a few hundred ms wide.
|
| 76 |
-
|
| 77 |
-
### Stage 1 — Fisher validation, teacher-forced
|
| 78 |
-
|
| 79 |
-
93,042 supervised frames, 1,213 positive (1.30%).
|
| 80 |
-
|
| 81 |
-
| Average Precision | ROC-AUC | mean p on positives | mean p on negatives |
|
| 82 |
-
|---|---|---|---|
|
| 83 |
-
| 0.352 | 0.928 | 0.486 | 0.182 |
|
| 84 |
-
|
| 85 |
-
| Threshold | Precision | Recall | F1@400ms | Rate vs. human |
|
| 86 |
-
|---|---|---|---|---|
|
| 87 |
-
| 0.70 | 0.206 | 0.305 | 0.246 | 1.48x |
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
### Stage 2 — TurnBench dev, out of domain
|
| 91 |
-
|
| 92 |
-
12 conversations / 24 dialogue passes, 6,553 supervised frames, 97 positive, 35 reference events,
|
| 93 |
-
human rate 1.52 events/min. Run on a **-23 LUFS loudness-normalised** copy: TurnBench ships 19.4 dB
|
| 94 |
-
apart internally and ~15 dB below Fisher, which moves where the head fires.
|
| 95 |
-
|
| 96 |
-
Two conditions: `keep` leaves the recording untouched (the agent's own backchannels stay in its
|
| 97 |
-
context, so the head can lean on them); `remove` erases them, which is the condition that matches
|
| 98 |
-
self-play and the render harness.
|
| 99 |
-
|
| 100 |
-
| Condition | AP | ROC-AUC | Best lag | Rate-matched thr | P | R | **F1** | Rate |
|
| 101 |
-
|---|---|---|---|---|---|---|---|---|
|
| 102 |
-
| `keep` | 0.147 | 0.810 | -1 | 0.650 | 0.222 | 0.229 | **0.225** | 1.03x |
|
| 103 |
-
| `remove` | 0.123 | 0.805 | -4 | 0.650 | 0.485 | 0.457 | **0.471** | 0.94x |
|
| 104 |
-
|
| 105 |
-
At best-F1 rather than rate-matched, `remove` reaches **F1 0.475** (thr 0.675, P 0.583 / R 0.400) but
|
| 106 |
-
sits at 0.69x the human rate — it buys precision by firing less.
|
| 107 |
-
|
| 108 |
-
Two caveats worth carrying:
|
| 109 |
-
|
| 110 |
-
* **The in-domain to out-of-domain drop is large and not checkpoint-specific**: AP 0.35 -> 0.12-0.15,
|
| 111 |
-
ROC-AUC 0.93 -> 0.81.
|
| 112 |
-
* **v14b's optimal lag on `remove` is -4 frames** (320 ms *before* the human onset), where the other
|
| 113 |
-
arms sit at -1/-2. AP at that lag is 0.231, against 0.123 at lag 0.
|
| 114 |
-
|
| 115 |
-
### Stage 3 — self-play obedience
|
| 116 |
-
|
| 117 |
-
100 dialogues x 80 s windows (133 min) in the partner-backchannelling condition. Thresholds are
|
| 118 |
-
calibrated **on generated streams**, never reused from validation — confidence is systematically
|
| 119 |
-
different in free generation than under teacher forcing.
|
| 120 |
-
|
| 121 |
-
| Condition | Threshold | Fires/min | **Obeyed** | **Realised** | clean bc | overrun | other | stranded |
|
| 122 |
-
|---|---|---|---|---|---|---|---|---|
|
| 123 |
-
| off | — | 0.00 | — | — | — | — | — | — |
|
| 124 |
-
| low | 0.740 | 1.03 | **100%** | **77%** | 106 | 31 | 0 | 0 |
|
| 125 |
-
| normal | 0.695 | 1.48 | **98%** | **69%** | 137 | 58 | 3 | 0 |
|
| 126 |
-
| high | 0.615 | 3.59 | **95%** | **65%** | 311 | 142 | 24 | 2 |
|
| 127 |
-
| max | 0.485 | 8.31 | **89%** | **65%** | 717 | 270 | 84 | 37 |
|
| 128 |
-
| *Partner (within-dialogue baseline)* | — | *1.51* | — | — | — | — | — | — |
|
| 129 |
-
|
| 130 |
-
* **`obeyed`** = the model started *something* at the trigger. **`realised`** = what it said was a
|
| 131 |
-
backchannel and stayed one. **`overrun`** = it began with a backchannel and then kept talking into
|
| 132 |
-
a turn. **`other`** = not a backchannel at all.
|
| 133 |
-
* At `normal`, only **3 of 198** triggers produced a non-backchannel.
|
| 134 |
-
* Obedience degrades gracefully: still 89% obeyed when driven to 5.5x the human rate.
|
| 135 |
-
* **Realised forms** are the expected inventory: `yeah` (~500), `right`, `okay`, `uh-huh`, `mm`,
|
| 136 |
-
`mhm`.
|
| 137 |
-
|
| 138 |
-
#### Controllability of `off`
|
| 139 |
-
|
| 140 |
-
With the head off, **zero** forced events fire, but the model still emits **0.70 spontaneous
|
| 141 |
-
lexical backchannels/min**, against 1.51/min for its partner in the same calls. So `off` means "no
|
| 142 |
-
head-driven backchannels", not silence — the purge lever moves this number but does not zero it.
|
| 143 |
-
|
| 144 |
-
#### Overrun
|
| 145 |
-
|
| 146 |
-
29% of obeyed triggers at `normal` (58/198) overran into a turn. The forced-decode harness measures
|
| 147 |
-
the same phenomenon far lower — 22/177 = 12% continued past a complete backchannel, of which 9
|
| 148 |
-
(5.1%) were true turn attempts — because its gate requires 3.2 s of clearance before firing, so it
|
| 149 |
-
only fires where a turn *cannot* start. Self-play lacks that constraint, so **treat 29% as the
|
| 150 |
-
pessimistic bound and 5% as what a clearance-gated deployment sees.**
|
| 151 |
-
|
| 152 |
-
### Choosing among the v14 arms
|
| 153 |
-
|
| 154 |
-
Rate-matched F1 on TurnBench, and obedience at a matched ~1.5 fires/min:
|
| 155 |
-
|
| 156 |
-
| ckpt | augmentation | Fisher F1 | TB `keep` F1 | TB `remove` F1 | obeyed | realised |
|
| 157 |
-
|---|---|---|---|---|---|---|
|
| 158 |
-
| v14a | none | .252 | **.267** | .348 | 87% | 51% |
|
| 159 |
-
| **v14b** | **deletions only** | .246 | .225 | .471 | **98%** | **69%** |
|
| 160 |
-
| v14c | insertions only | .244 | **.270** | .324 | 93% | 63% |
|
| 161 |
-
| v14d | both | .246 | .257 | **.479** | 93% | 53% |
|
| 162 |
-
|
| 163 |
-
1. **Fisher separates nothing** — all four within 0.008 F1 and 0.002 AUC on 1,213 positive frames.
|
| 164 |
-
2. **`remove` splits them into two pairs**: v14b/v14d at .471/.479 against v14a/v14c at .348/.324, a
|
| 165 |
-
~45% relative gap driven by precision. But v14b vs v14d is a coin flip on 35 reference events.
|
| 166 |
-
3. **`keep` ranks them differently** and puts v14b last. `keep` is the easier condition and the one
|
| 167 |
-
that does not match how the model is actually run.
|
| 168 |
-
4. **Obedience is where they separate, and v14b wins outright** — at every floor, on both metrics,
|
| 169 |
-
with 3 non-backchannels out of 198 against 26 for v14a.
|
| 170 |
-
|
| 171 |
-
Counterintuitive and worth restating: **deletions-only beats insertions-only and both**, even though
|
| 172 |
-
insertion is the arm meant to teach the model to obey a backchannel request.
|
| 173 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
- backchannel
|
| 16 |
---
|
| 17 |
|
|
|
|
| 18 |
|
| 19 |
+
# PersonaPlex Backchannel Head
|
| 20 |
|
| 21 |
+
This HF repository contains a LoRA adapter for [PersonaPlex-7B](https://huggingface.co/nvidia/personaplex-7b-v1) that adds a lightweight, controllable backchannel head, introduced in [Controlling Backchannels in Streamable Full-Duplex Models](https://arxiv.org/abs/2609.29418).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
|
| 23 |
+
About our work: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.
|
|
|
|
|
|
|
|
|
|
| 24 |
|
| 25 |
+
Please refer to the [codebase](https://github.com/MaikeZuefle/bcmore) for the usage of the model.
|
| 26 |
|
| 27 |
+
For more information, please have a look at the paper.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
+
## Citation
|
| 30 |
|
| 31 |
+
If you use this model, please cite:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
+
```bibtex
|
| 34 |
+
@misc{züfle2026controllingbackchannelsstreamablefullduplex,
|
| 35 |
+
title={Controlling Backchannels in Streamable Full-duplex Models},
|
| 36 |
+
author={Maike Züfle and Peter Polák and Sefik Emre Eskimez and Jan Niehues and Peter Bell and Ondřej Klejch},
|
| 37 |
+
year={2026},
|
| 38 |
+
eprint={2609.29418},
|
| 39 |
+
archivePrefix={arXiv},
|
| 40 |
+
primaryClass={cs.CL},
|
| 41 |
+
url={https://arxiv.org/abs/2609.29418},
|
| 42 |
+
}
|
| 43 |
+
```
|