Skip to main navigation Skip to search Skip to main content

Referring Image Segmentation Using Text Supervision

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to localize the target object. Hence, we propose a novel weakly-supervised RIS framework to formulate the target localization problem as a classification process to differentiate between positive and negative text expressions. While the referring text expressions for an image are used as positive expressions, the referring text expressions from other images can be used as negative expressions for this image. Our framework has three main novelties. First, we propose a bilateral prompt method to facilitate the classification process, by harmonizing the domain discrepancy between visual and linguistic features. Second, we propose a calibration method to reduce noisy background information and improve the correctness of the response maps for target object localization. Third, we propose a positive response map selection strategy to generate high-quality pseudo-labels from the enhanced response maps, for training a segmentation network for RIS inference. For evaluation, we propose a new metric to measure localization accuracy. Experiments on four benchmarks show that our framework achieves promising performances to existing fully-supervised RIS methods while outperforming state-of-the-art weakly-supervised methods adapted from related areas. Code is available at https://github.com/fawnliu/TRIS. ©2023 IEEE.
Original languageEnglish
Title of host publicationProceedings - 2023 IEEE/CVF International Conference on Computer Vision ICCV 2023
PublisherIEEE
Pages22067-22077
ISBN (Electronic)979-8-3503-0718-4
ISBN (Print)979-8-3503-0719-1
DOIs
Publication statusPublished - Oct 2023
Event2023 IEEE International Conference on Computer Vision (ICCV 2023) - Paris Convention Center , Paris, France
Duration: 2 Oct 20236 Oct 2023
https://iccv2023.thecvf.com/

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2023 IEEE International Conference on Computer Vision (ICCV 2023)
Abbreviated titleICCV23
PlaceFrance
CityParis
Period2/10/236/10/23
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Fingerprint

Dive into the research topics of 'Referring Image Segmentation Using Text Supervision'. Together they form a unique fingerprint.

Cite this