Abstract
This paper tackles a recently proposed Video Corpus Moment Retrieval task. This task is essential because advanced video retrieval applications should enable users to retrieve a precise moment from a large video corpus. We propose a novel CONtextual QUery-awarE Ranking∼(CONQUER) model for effective moment localization and ranking. CONQUER explores query context for multi-modal fusion and representation learning in two different steps. The first step derives fusion weights for the adaptive combination of multi-modal video content. The second step performs bi-directional attention to tightly couple video and query as a single joint representation for moment localization. As query context is fully engaged in video representation learning, from feature fusion to transformation, the resulting feature is user-centered and has a larger capacity in capturing multi-modal signals specific to query. We conduct studies on two datasets, TVR for closed-world TV episodes and DiDeMo for open-world user-generated videos, to investigate the potential advantages of fusing video and query online as a joint representation for moment retrieval.
| Original language | English |
|---|---|
| Title of host publication | MM' 21 |
| Subtitle of host publication | Proceedings of the 29th ACM International Conference on Multimedia |
| Place of Publication | New York |
| Publisher | Association for Computing Machinery |
| Pages | 3900-3908 |
| Number of pages | 9 |
| ISBN (Print) | 978-1-4503-8651-7 |
| DOIs | |
| Publication status | Published - 2021 |
| Event | 29th ACM International Conference on Multimedia (MM 2021) - Hybrid, Chengdu, China Duration: 20 Oct 2021 → 24 Oct 2021 https://2021.acmmm.org/ |
Publication series
| Name | MM 2021 - Proceedings of the 29th ACM International Conference on Multimedia |
|---|
Conference
| Conference | 29th ACM International Conference on Multimedia (MM 2021) |
|---|---|
| Abbreviated title | MM '21 |
| Place | China |
| City | Chengdu |
| Period | 20/10/21 → 24/10/21 |
| Internet address |
Bibliographical note
Research Unit(s) information for this publication is provided by the author(s) concerned.Research Keywords
- cross-modal retrieval
- moment localization with natural language
RGC Funding Information
- RGC-funded
Fingerprint
Dive into the research topics of 'CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver