Skip to main navigation Skip to search Skip to main content

ZOOM: Learning Video Mirror Detection with Extremely-Weak Supervision

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Mirror detection is an active research topic in computer vision. However, all existing mirror detectors learn mirror representations from large-scale pixel-wise datasets, which are tedious and expensive to obtain. Although weakly-supervised learning has been widely explored in related topics, we note that popular weak supervision signals (e.g., bounding boxes, scribbles, points) still require some efforts from the user to locate the target objects, with a strong assumption that the images to annotate always contain the target objects. Such an assumption may result in the over-segmentation of mirrors. Our key idea of this work is that the existence of mirrors over a time period may serve as a weak supervision to train a mirror detector, for two reasons. First, if a network can predict the existence of mirrors, it can essentially locate the mirrors. Second, we observe that the reflected contents of a mirror tend to be similar to those in adjacent frames, but exhibit considerable contrast to regions in far-away frames (e.g., non-mirror frames). In this paper, we propose ZOOM, the first method to learn robust mirror representations from extremely weak annotations of per-frame ZerO-One Mirror indicators in videos. The key insight of ZOOM is to model the similarity and contrast (between the mirror and non-mirror regions) in temporal variations to locate and segment the mirrors. To this end, we propose a novel fusion strategy to leverage temporal consistency information for mirror localization and a novel temporal similarity-contrast modeling module for mirror segmentation. We construct a new video mirror dataset for training and evaluation. Experimental results under new and standard metrics show that ZOOM performs favorably against existing fully-supervised mirror detection methods. Copyright © 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
Original languageEnglish
Title of host publicationProceedings of the 38th AAAI Conference on Artificial Intelligence
EditorsMichael Wooldridge, Jennifer Dy, Sriraam Natarajan
PublisherAAAI Press
Pages6315-6323
ISBN (Print)1-57735-887-2, 978-1-57735-887-9
DOIs
Publication statusPublished - 2024
Event38th AAAI Conference on Artificial Intelligence (AAAI 2024) - Vancouver, Canada
Duration: 20 Feb 202427 Feb 2024

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
PublisherAssociation for the Advancement of Artificial Intelligence
Number6
Volume38
ISSN (Print)2159-5399
ISSN (Electronic)2374-3468

Conference

Conference38th AAAI Conference on Artificial Intelligence (AAAI 2024)
Abbreviated titleAAAI-24
PlaceCanada
CityVancouver
Period20/02/2427/02/24

Funding

This project is in part supported by a GRF grant from the Research Grants Council of Hong Kong (No.: 11211223) and an SRG grant from City University of Hong Kong (No.: 7005674).

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'ZOOM: Learning Video Mirror Detection with Extremely-Weak Supervision'. Together they form a unique fingerprint.

Cite this