Skip to main navigation Skip to search Skip to main content

ZA-SLAM: Leveraging Vision-Language Model for Zero-Shot Acoustic SLAM

  • Zhuochen Yu
  • , David K.Y. Yau
  • , Yijie Shen
  • , Xiaoran Fan
  • , Tao Chen
  • , Qun Song*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

2 Downloads (CityUHK Scholars)

Abstract

Existing acoustic indoor location sensing systems are limited by the need for extensive data collection and model retraining in unseen environments. This paper introduces ZA-SLAM, a novel zero-shot acoustic Simultaneous Localization and Mapping (SLAM) system that can be deployed in unseen environments without model retraining. Our core idea is to train an acoustic encoder that inherits the generalization capabilities of pre-trained Vision-Language Models (VLMs), which show superiority in tasks like zero-shot visual SLAM. To achieve this goal, we perform Acoustic-Visual Feature Alignment to enable the acoustic encoder to generate features aligned with visual features from VLMs. To select high-quality images for effective alignment, we design a Semantic-Guided Image Selection that filters out low-quality collected images caused by factors like abrupt view changes, occlusions, and uninformative views. Furthermore, we address the challenge of false positive loop closures in structurally similar locations with the Learning-Based Trajectory Reachability Matching that validates loop closures leveraging IMU trajectory features. Extensive real-world experiments demonstrate that our system achieves comparable SLAM performance to retraining-based acoustic SLAM, and much improved performance compared to existing zero-shot Wi-Fi and geomagnetic SLAM systems. Our system achieves a mean mapping error of 0.56 m and a localization error of 0.78 m across multiple unseen environments.
© 2026 Copyright held by the owner/author(s).
Original languageEnglish
Title of host publicationMobiSys '26
Subtitle of host publicationProceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services
PublisherAssociation for Computing Machinery
Pages492-505
ISBN (Print)979-8-4007-2027-7
DOIs
Publication statusPublished - 20 Jun 2026
Event24th ACM International Conference on Mobile Systems, Applications, and Services (MobiSys 2026) - Cambridge, United Kingdom
Duration: 21 Jun 202625 Jun 2026
https://www.sigmobile.org/mobisys/2026/

Conference

Conference24th ACM International Conference on Mobile Systems, Applications, and Services (MobiSys 2026)
Abbreviated titleMobiSys 2026
PlaceUnited Kingdom
CityCambridge
Period21/06/2625/06/26
Internet address

Bibliographical note

Research Unit(s) information for this publication is provided by the author(s) concerned.

Research Keywords

  • Simultaneous Localization and Mapping
  • Acoustic Sensing
  • Zero-Shot Learning
  • Vision-Language Models

Publisher's Copyright Statement

  • This full text is made available under CC-BY-NC-ND 4.0. https://creativecommons.org/licenses/by-nc-nd/4.0/

Fingerprint

Dive into the research topics of 'ZA-SLAM: Leveraging Vision-Language Model for Zero-Shot Acoustic SLAM'. Together they form a unique fingerprint.

Cite this