Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM

Jia Wan*, Qiangqiang Wu, Wei Lin, Antoni Chan

*Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

5 Citations (Scopus)

Abstract

The existing crowd counting models require extensive training data, which is time-consuming to annotate. To tackle this issue, we propose a simple yet effective crowd counting method by utilizing the Segment-Everything-Everywhere Model (SEEM), an adaptation of the Segmentation Anything Model (SAM), to generate pseudo-labels for training crowd counting models. However, our initial investigation reveals that SEEM’s performance in dense crowd scenes is limited, primarily due to the omission of many persons in high-density areas. To overcome this limitation, we propose an adaptive resolution SEEM to handle the scale variations, occlusions, and overlapping of people within crowd scenes. Alongside this, we introduce a robust localization method, based on Gaussian Mixture Models, for predicting the head positions in the predicted people masks. Given the mask and point pseudo-labels, we propose a robust loss function, which is designed to exclude uncertain regions based on SEEM’s predictions, thereby enhancing the training process of the counting network. Finally, we propose an iterative method for generating pseudo-labels. This method aims at improving the quality of the segmentation masks by identifying more tiny persons in high-density regions, which are often missed in the first pseudo-labeling iteration. Overall, our proposed method achieves the best unsupervised performance in crowd counting, while also being comparable to some classic supervised fully methods. This makes it a highly effective and versatile tool for crowd counting, especially in situations where labeled data is not available. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024
Subtitle of host publication18th European Conference, Proceedings, Part LVII
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer, Cham
Pages478-495
Edition1
ISBN (Electronic)978-3-031-72998-0
ISBN (Print)978-3-031-72997-3
DOIs
Publication statusPublished - 2024
Event18th European Conference on Computer Vision (ECCV 2024) - MiCo Milano, Milan, Italy
Duration: 29 Sept 20244 Oct 2024
https://eccv.ecva.net/

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15115 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision (ECCV 2024)
Abbreviated titleECCV2024
PlaceItaly
CityMilan
Period29/09/244/10/24
Internet address

Funding

This work was supported by a Strategic Research Grant from City University of Hong Kong (Project No. 7005665).

Research Keywords

  • Crowd Counting
  • Crowd Localization
  • Segment Anything

Fingerprint

Dive into the research topics of 'Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM'. Together they form a unique fingerprint.

Cite this