Skip to main navigation Skip to search Skip to main content

CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

This paper tackles a recently proposed Video Corpus Moment Retrieval task. This task is essential because advanced video retrieval applications should enable users to retrieve a precise moment from a large video corpus. We propose a novel CONtextual QUery-awarE Ranking∼(CONQUER) model for effective moment localization and ranking. CONQUER explores query context for multi-modal fusion and representation learning in two different steps. The first step derives fusion weights for the adaptive combination of multi-modal video content. The second step performs bi-directional attention to tightly couple video and query as a single joint representation for moment localization. As query context is fully engaged in video representation learning, from feature fusion to transformation, the resulting feature is user-centered and has a larger capacity in capturing multi-modal signals specific to query. We conduct studies on two datasets, TVR for closed-world TV episodes and DiDeMo for open-world user-generated videos, to investigate the potential advantages of fusing video and query online as a joint representation for moment retrieval.
Original languageEnglish
Title of host publicationMM' 21
Subtitle of host publicationProceedings of the 29th ACM International Conference on Multimedia
Place of PublicationNew York
PublisherAssociation for Computing Machinery
Pages3900-3908
Number of pages9
ISBN (Print)978-1-4503-8651-7
DOIs
Publication statusPublished - 2021
Event29th ACM International Conference on Multimedia (MM 2021) - Hybrid, Chengdu, China
Duration: 20 Oct 202124 Oct 2021
https://2021.acmmm.org/

Publication series

NameMM 2021 - Proceedings of the 29th ACM International Conference on Multimedia

Conference

Conference29th ACM International Conference on Multimedia (MM 2021)
Abbreviated titleMM '21
PlaceChina
CityChengdu
Period20/10/2124/10/21
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Research Keywords

  • cross-modal retrieval
  • moment localization with natural language

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval'. Together they form a unique fingerprint.

Cite this