Instructions to use JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head with PEFT:
Task type is invalid.
- Moshi
How to use JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head with Moshi:
# pip install moshi # Run the interactive web server python -m moshi.server --hf-repo "JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head" # Then open https://localhost:8998 in your browser
# pip install moshi import torch from moshi.models import loaders # Load checkpoint info from HuggingFace checkpoint = loaders.CheckpointInfo.from_hf_repo("JSALT2026-Conv-AI-Simulator/personaplex-fisher-bc-head") # Load the Mimi audio codec mimi = checkpoint.get_mimi(device="cuda") mimi.set_num_codebooks(8) # Encode audio (24kHz, mono) wav = torch.randn(1, 1, 24000 * 10) # [batch, channels, samples] with torch.no_grad(): codes = mimi.encode(wav.cuda()) decoded = mimi.decode(codes) - Notebooks
- Google Colab
- Kaggle
Add model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,173 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: nvidia/personaplex-7b-v1
|
| 3 |
+
library_name: peft
|
| 4 |
+
license: other
|
| 5 |
+
license_name: nvidia-open-model-license
|
| 6 |
+
license_link: https://huggingface.co/nvidia/personaplex-7b-v1
|
| 7 |
+
datasets:
|
| 8 |
+
- JSALT2026-Conv-AI-Simulator/fisher-v1
|
| 9 |
+
tags:
|
| 10 |
+
- lora
|
| 11 |
+
- personaplex
|
| 12 |
+
- moshi
|
| 13 |
+
- full-duplex
|
| 14 |
+
- speech-to-speech
|
| 15 |
+
- backchannel
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
# personaplex-fisher-lora-backchannel-v14b
|
| 19 |
+
|
| 20 |
+
### Augmentation recipe
|
| 21 |
+
|
| 22 |
+
| lever | value | effect |
|
| 23 |
+
|---|---|---|
|
| 24 |
+
| `--backchannel_aug_prob` | 0.3 | 30% of examples are augmented |
|
| 25 |
+
| `--backchannel_aug_delete_share` | 1.0 | every augmented example is a *deletion* example β no insertions |
|
| 26 |
+
| `--backchannel_aug_delete_prob` | 0.5 | half the backchannels in such an example are removed |
|
| 27 |
+
| `--backchannel_aug_purge_prob` | 0.05 | 5% of examples have **every** backchannel erased, drawn independently of `aug_prob` |
|
| 28 |
+
|
| 29 |
+
The purge is the important one. At `delete_prob 0.5` alone, every augmented example still contains
|
| 30 |
+
backchannels, so the model can learn "fewer here" but never "none" β and "none" is the only signal
|
| 31 |
+
that acts on *spontaneous* production. Purged examples keep their head positives, so head
|
| 32 |
+
supervision is identical across all four v14 arms and the purge acts only on the LM.
|
| 33 |
+
|
| 34 |
+
## Training
|
| 35 |
+
|
| 36 |
+
* **Data:** [`JSALT2026-Conv-AI-Simulator/fisher-v1`](https://huggingface.co/datasets/JSALT2026-Conv-AI-Simulator/fisher-v1)
|
| 37 |
+
(Fisher telephone conversations), `max_length` 2000 frames (~160 s), 4 epochs, batch 1 x
|
| 38 |
+
grad-accum 16.
|
| 39 |
+
* **Backchannel labels** are derived from word alignments (Fisher's own
|
| 40 |
+
`utterances[].backchannels` column is empty): each speaker is split into inter-pausal units, and
|
| 41 |
+
one is marked a backchannel when it is short (`max_duration 1.5`), passes a content gate, and did
|
| 42 |
+
not take the floor (`floor_gap 2.0`, retaining floor status across gaps up to 2 s).
|
| 43 |
+
`--no-backchannel_include_laughter` β unlike `v2`, listener laughter does **not** count.
|
| 44 |
+
`--backchannel_turn_clearance 1.5`, `--drop_non_speech_words`.
|
| 45 |
+
* Only frames where a backchannel is *possible* (agent silent, partner talking,
|
| 46 |
+
`agent_silence_margin 0.24`) are supervised; the rest are ignored rather than taught as negatives.
|
| 47 |
+
|
| 48 |
+
### Architecture
|
| 49 |
+
|
| 50 |
+
* **LoRA:** attention only, on the temporal transformer β
|
| 51 |
+
`^decoder\.model\.layers\.\d+\.(self_attn\.(q_proj|k_proj|v_proj|o_proj)\.linear)$` β
|
| 52 |
+
`r=16`, `alpha=32`, `dropout=0.05`.
|
| 53 |
+
* **Backchannel head:** `Linear(hidden -> 256) -> GELU -> Linear(256 -> 1)`, reading a single frame
|
| 54 |
+
of the trunk's hidden state. `context=0`, i.e. **no** temporal convolution β that is the `claims`
|
| 55 |
+
line of checkpoints.
|
| 56 |
+
* **Trainable tokens:** the `<bc>` marker is token id **4**, the `<0x00>` byte-fallback slot β the
|
| 57 |
+
one slot the tokenizer can never emit (it occurs zero times across all 11,396 Fisher rows). Its
|
| 58 |
+
embedding row is trained in both `decoder.model.embed_tokens` and
|
| 59 |
+
`depth_decoder.text_embed_tokens`; the rest of both tables stays frozen.
|
| 60 |
+
* **Loss:** focal (`gamma=2.0`, `alpha=0.9`), backchannel loss weight `1.0`, head bias initialised
|
| 61 |
+
to prior `0.016`, head LR `1e-3`, LM LR `2e-5`. Semantic codebook weight `1.0`, acoustic `0.02`,
|
| 62 |
+
pad-text loss weight `0.3`.
|
| 63 |
+
|
| 64 |
+
## Evaluation
|
| 65 |
+
|
| 66 |
+
Three measurements that answer three different questions:
|
| 67 |
+
|
| 68 |
+
| measurement | what it asks | material |
|
| 69 |
+
|---|---|---|
|
| 70 |
+
| Fisher head eval | can the head rank frames? | Fisher validation, teacher-forced |
|
| 71 |
+
| TurnBench head eval | does that survive out of domain, against human annotations? | TurnBench dev, teacher-forced |
|
| 72 |
+
| self-play obedience | when the head fires, does the model produce a backchannel? | Fisher-domain self-play |
|
| 73 |
+
|
| 74 |
+
Matching tolerance is 5 frames (400 ms) in both head evals β exact-frame agreement is too strict a
|
| 75 |
+
bar for a phenomenon whose natural placement window is a few hundred ms wide.
|
| 76 |
+
|
| 77 |
+
### Stage 1 β Fisher validation, teacher-forced
|
| 78 |
+
|
| 79 |
+
93,042 supervised frames, 1,213 positive (1.30%).
|
| 80 |
+
|
| 81 |
+
| Average Precision | ROC-AUC | mean p on positives | mean p on negatives |
|
| 82 |
+
|---|---|---|---|
|
| 83 |
+
| 0.352 | 0.928 | 0.486 | 0.182 |
|
| 84 |
+
|
| 85 |
+
| Threshold | Precision | Recall | F1@400ms | Rate vs. human |
|
| 86 |
+
|---|---|---|---|---|
|
| 87 |
+
| 0.70 | 0.206 | 0.305 | 0.246 | 1.48x |
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
### Stage 2 β TurnBench dev, out of domain
|
| 91 |
+
|
| 92 |
+
12 conversations / 24 dialogue passes, 6,553 supervised frames, 97 positive, 35 reference events,
|
| 93 |
+
human rate 1.52 events/min. Run on a **-23 LUFS loudness-normalised** copy: TurnBench ships 19.4 dB
|
| 94 |
+
apart internally and ~15 dB below Fisher, which moves where the head fires.
|
| 95 |
+
|
| 96 |
+
Two conditions: `keep` leaves the recording untouched (the agent's own backchannels stay in its
|
| 97 |
+
context, so the head can lean on them); `remove` erases them, which is the condition that matches
|
| 98 |
+
self-play and the render harness.
|
| 99 |
+
|
| 100 |
+
| Condition | AP | ROC-AUC | Best lag | Rate-matched thr | P | R | **F1** | Rate |
|
| 101 |
+
|---|---|---|---|---|---|---|---|---|
|
| 102 |
+
| `keep` | 0.147 | 0.810 | -1 | 0.650 | 0.222 | 0.229 | **0.225** | 1.03x |
|
| 103 |
+
| `remove` | 0.123 | 0.805 | -4 | 0.650 | 0.485 | 0.457 | **0.471** | 0.94x |
|
| 104 |
+
|
| 105 |
+
At best-F1 rather than rate-matched, `remove` reaches **F1 0.475** (thr 0.675, P 0.583 / R 0.400) but
|
| 106 |
+
sits at 0.69x the human rate β it buys precision by firing less.
|
| 107 |
+
|
| 108 |
+
Two caveats worth carrying:
|
| 109 |
+
|
| 110 |
+
* **The in-domain to out-of-domain drop is large and not checkpoint-specific**: AP 0.35 -> 0.12-0.15,
|
| 111 |
+
ROC-AUC 0.93 -> 0.81.
|
| 112 |
+
* **v14b's optimal lag on `remove` is -4 frames** (320 ms *before* the human onset), where the other
|
| 113 |
+
arms sit at -1/-2. AP at that lag is 0.231, against 0.123 at lag 0.
|
| 114 |
+
|
| 115 |
+
### Stage 3 β self-play obedience
|
| 116 |
+
|
| 117 |
+
100 dialogues x 80 s windows (133 min) in the partner-backchannelling condition. Thresholds are
|
| 118 |
+
calibrated **on generated streams**, never reused from validation β confidence is systematically
|
| 119 |
+
different in free generation than under teacher forcing.
|
| 120 |
+
|
| 121 |
+
| Condition | Threshold | Fires/min | **Obeyed** | **Realised** | clean bc | overrun | other | stranded |
|
| 122 |
+
|---|---|---|---|---|---|---|---|---|
|
| 123 |
+
| off | β | 0.00 | β | β | β | β | β | β |
|
| 124 |
+
| low | 0.740 | 1.03 | **100%** | **77%** | 106 | 31 | 0 | 0 |
|
| 125 |
+
| normal | 0.695 | 1.48 | **98%** | **69%** | 137 | 58 | 3 | 0 |
|
| 126 |
+
| high | 0.615 | 3.59 | **95%** | **65%** | 311 | 142 | 24 | 2 |
|
| 127 |
+
| max | 0.485 | 8.31 | **89%** | **65%** | 717 | 270 | 84 | 37 |
|
| 128 |
+
| *Partner (within-dialogue baseline)* | β | *1.51* | β | β | β | β | β | β |
|
| 129 |
+
|
| 130 |
+
* **`obeyed`** = the model started *something* at the trigger. **`realised`** = what it said was a
|
| 131 |
+
backchannel and stayed one. **`overrun`** = it began with a backchannel and then kept talking into
|
| 132 |
+
a turn. **`other`** = not a backchannel at all.
|
| 133 |
+
* At `normal`, only **3 of 198** triggers produced a non-backchannel.
|
| 134 |
+
* Obedience degrades gracefully: still 89% obeyed when driven to 5.5x the human rate.
|
| 135 |
+
* **Realised forms** are the expected inventory: `yeah` (~500), `right`, `okay`, `uh-huh`, `mm`,
|
| 136 |
+
`mhm`.
|
| 137 |
+
|
| 138 |
+
#### Controllability of `off`
|
| 139 |
+
|
| 140 |
+
With the head off, **zero** forced events fire, but the model still emits **0.70 spontaneous
|
| 141 |
+
lexical backchannels/min**, against 1.51/min for its partner in the same calls. So `off` means "no
|
| 142 |
+
head-driven backchannels", not silence β the purge lever moves this number but does not zero it.
|
| 143 |
+
|
| 144 |
+
#### Overrun
|
| 145 |
+
|
| 146 |
+
29% of obeyed triggers at `normal` (58/198) overran into a turn. The forced-decode harness measures
|
| 147 |
+
the same phenomenon far lower β 22/177 = 12% continued past a complete backchannel, of which 9
|
| 148 |
+
(5.1%) were true turn attempts β because its gate requires 3.2 s of clearance before firing, so it
|
| 149 |
+
only fires where a turn *cannot* start. Self-play lacks that constraint, so **treat 29% as the
|
| 150 |
+
pessimistic bound and 5% as what a clearance-gated deployment sees.**
|
| 151 |
+
|
| 152 |
+
### Choosing among the v14 arms
|
| 153 |
+
|
| 154 |
+
Rate-matched F1 on TurnBench, and obedience at a matched ~1.5 fires/min:
|
| 155 |
+
|
| 156 |
+
| ckpt | augmentation | Fisher F1 | TB `keep` F1 | TB `remove` F1 | obeyed | realised |
|
| 157 |
+
|---|---|---|---|---|---|---|
|
| 158 |
+
| v14a | none | .252 | **.267** | .348 | 87% | 51% |
|
| 159 |
+
| **v14b** | **deletions only** | .246 | .225 | .471 | **98%** | **69%** |
|
| 160 |
+
| v14c | insertions only | .244 | **.270** | .324 | 93% | 63% |
|
| 161 |
+
| v14d | both | .246 | .257 | **.479** | 93% | 53% |
|
| 162 |
+
|
| 163 |
+
1. **Fisher separates nothing** β all four within 0.008 F1 and 0.002 AUC on 1,213 positive frames.
|
| 164 |
+
2. **`remove` splits them into two pairs**: v14b/v14d at .471/.479 against v14a/v14c at .348/.324, a
|
| 165 |
+
~45% relative gap driven by precision. But v14b vs v14d is a coin flip on 35 reference events.
|
| 166 |
+
3. **`keep` ranks them differently** and puts v14b last. `keep` is the easier condition and the one
|
| 167 |
+
that does not match how the model is actually run.
|
| 168 |
+
4. **Obedience is where they separate, and v14b wins outright** β at every floor, on both metrics,
|
| 169 |
+
with 3 non-backchannels out of 198 against 26 for v14a.
|
| 170 |
+
|
| 171 |
+
Counterintuitive and worth restating: **deletions-only beats insertions-only and both**, even though
|
| 172 |
+
insertion is the arm meant to teach the model to obey a backchannel request.
|
| 173 |
+
|