Abstract
Predicting long-term scanpaths in panoramic videos requires effective fusion of multimodal inputs, including visual content and historical gaze sequences. Existing methods typically process these modalities independently, ignoring their intrinsic temporal-spatial correlations and thereby limiting the fidelity of multimodal distribution modeling. This paper introduces a unified framework that explicitly aligns visual and gaze modalities in both time and space. A hybrid-granularity Transformer is proposed to jointly encode global semantic structures and local fine-grained dynamics, enabling more accurate long-term dependency modeling. To further enhance cross-modal fusion, a contrastive learning strategy is employed to improve the alignment of visual features and scanpath representations. For scanpath generation, a lightweight optimization-based sampler-guided by a physics-inspired proxy viewer-is integrated to produce smooth and realistic gaze trajectories without relying on heuristic sampling. Evaluations on VRW-23 and CVPR-18 demonstrate consistent state-of-the-art performance, confirming the effectiveness of the proposed temporal-spatial multimodal alignment and hybrid-granularity architecture. © 2026 Elsevier Ltd.
| Original language | English |
|---|---|
| Article number | 113480 |
| Journal | Pattern Recognition |
| Volume | 179, Part A |
| Online published | 15 Mar 2026 |
| DOIs | |
| Publication status | Online published - 15 Mar 2026 |
Funding
This paper is supported in part by the National Natural Science Foundation of China under No. 62472124 and Shenzhen Colleges and Universities Stable Support Program under No. GXWD20220811170130002.
Research Keywords
- Panoramic video
- Scanpath prediction
- Multimodal learning
- Contrastive learning
- Transformer architecture
Fingerprint
Dive into the research topics of 'Gradient descent-driven sampling for multimodal long-term scanpath prediction in panoramic videos'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver