THU-IAR/MIntRec2.0
Updated • 2.9k • 2
our model is first compact (2.2B) emotion intelligence Multimodel Langugage Model. It integrates a pretrained LLM with modality-specific encoders and experts-based fusion encoder to handle a broad spectrum of affective tasks in one model.
Nano-EmoX supports six core tasks within one model: 1), Multimodal Sentiment Analysis. 2), Multimodal Emotion Recognition. 3), Open-Vocabulary Multimodal Emotion Recognition. 4), Multimodal Intention Recognition. 5), Emotion Reason Inference. 6), Empathic Response Generation
If this work has been helpful or inspiring to your research, please consider cite our article:
BibTeX:
@InProceedings{Huang_2026_CVPR,
author = {Huang, Jiahao and Lin, Fengyan and Yang, Xuechao and Feng, Chen and Zhu, Kexin and Yang, Xu and Chen, Zhide},
title = {Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {22986-22997}
}