PEFT
Safetensors
Moshi
lora
personaplex
full-duplex
speech-to-speech
backchannel
maikezu's picture
Update README.md
2e6ab4d verified
|
Raw History Blame Contribute Delete
2 kB
metadata
base_model: nvidia/personaplex-7b-v1
library_name: peft
license: other
license_name: nvidia-open-model-license
license_link: https://huggingface.co/nvidia/personaplex-7b-v1
datasets:
  - JSALT2026-Conv-AI-Simulator/fisher-v1
tags:
  - lora
  - personaplex
  - moshi
  - full-duplex
  - speech-to-speech
  - backchannel

PersonaPlex Backchannel Head

This HF repository contains a LoRA adapter for PersonaPlex-7B that adds a lightweight, controllable backchannel head, introduced in Controlling Backchannels in Streamable Full-Duplex Models.

About our work: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.

Please refer to the codebase for the usage of the model.

For more information, please have a look at the paper.

Citation

If you use this model, please cite:

@misc{züfle2026controllingbackchannelsstreamablefullduplex,
      title={Controlling Backchannels in Streamable Full-duplex Models}, 
      author={Maike Züfle and Peter Polák and Sefik Emre Eskimez and Jan Niehues and Peter Bell and Ondřej Klejch},
      year={2026},
      eprint={2609.29418},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.29418}, 
}