Abstract
Multimodal emotion recognition in real-world conversations often suffers from noise caused by information redundancy and erroneous correlational dependencies. These two issues are further amplified when directly using off-the-shelf large pre-trained models. Such models often hallucinate emotional cues, due to the insufficient modeling of structured dialogue relationships and inadequate tracking of emotion related states. To address these issues, this paper proposes a novel fusion framework. It integrates dialogue state tracking with large-scale model reasoning to trace key emotion-related evidence, thereby alleviating information redundancy. Using the traced evidence, it further employs a trainable dynamic graph learning mechanism to perform adaptive relational refinement, so as to deal with erroneous correlations. Specifically, we design a state token mechanism with large model guidance that serves as a dynamic emotion modeling strategy, continuously tracking and aggregating salient emotional cues across modalities and turns. On top of this, a variational dynamic graph structure learning module is introduced to jointly refine intra- and inter-modal dependencies, suppressing redundant edges while reinforcing emotion-discriminative ones. In addition, to facilitate better adaptive tracking and refinement, we further incorporate a task-oriented curriculum learning scheme and an emotion-guided contrastive learning strategy, enabling better distinguishing subtle and conflicting emotional states. Finally, extensive experiments on three real-world datasets demonstrate effectiveness in a dynamic multimodal conversation setting with changing emotions. © 2026 Elsevier Ltd.
| Original language | English |
|---|---|
| Article number | 114018 |
| Number of pages | 10 |
| Journal | Pattern Recognition |
| Volume | 180 |
| Issue number | Part A |
| Online published | 21 May 2026 |
| DOIs | |
| Publication status | Online published - 21 May 2026 |
Research Keywords
- Dialogue state tracking
- Fusion framework
- Graph neural network
- Large model
- Multimodal emotion recognition
Fingerprint
Dive into the research topics of 'Multimodal emotion recognition via large model guided dialogue state tracking with dynamic graph refinement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver