EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
Abstract
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
Community
EmoRES improves emotional speech generation without retraining by decomposing emotion steering vectors into shared and emotion-specific components and selectively strengthening the latter, improving emotion control across different TTS backbones.
Code will be available at: https://github.com/facebookresearch/EmoRES-TTS
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows (2026)
- EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis (2026)
- Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS (2026)
- Controllable Affective Generation via Latent Vector Steering (2026)
- Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language (2026)
- EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation (2026)
- Toward Human-Aligned Judgement of Speech Emotion Similarity (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.38157 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper