Skip to main navigation Skip to search Skip to main content

Learning Tracking Representations from Single Point Annotations

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Existing deep trackers are typically trained with largescale video frames with annotated bounding boxes. However, these bounding boxes are expensive and timeconsuming to annotate, in particular for large scale datasets. In this paper, we propose to learn tracking representations from single point annotations (i.e., 4.5× faster to annotate than the traditional bounding box) in a weakly supervised manner. Specifically, we propose a soft contrastive learning (SoCL) framework that incorporates target objectness prior into end-to-end contrastive learning. Our SoCL consists of adaptive positive and negative sample generation, which is memory-efficient and effective for learning tracking representations. We apply the learned representation of SoCL to visual tracking and show that our method can 1) achieve better performance than the fully supervised baseline trained with box annotations under the same annotation time cost; 2) achieve comparable performance of the fully supervised baseline by using the same number of training frames and meanwhile reducing annotation time cost by 78% and total fees by 85%; 3) be robust to annotation noise. © 2024 IEEE
Original languageEnglish
Title of host publicationProceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops
Subtitle of host publicationCVPRW 2024
PublisherIEEE
Pages2606-2615
ISBN (Electronic)9798350365474
ISBN (Print)979-8-3503-6548-1
DOIs
Publication statusPublished - 2024
Event2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW 2024) - Seattle, United States
Duration: 16 Jun 202422 Jun 2024

Publication series

NameIEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops
ISSN (Print)2160-7508
ISSN (Electronic)2160-7516

Conference

Conference2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW 2024)
PlaceUnited States
CitySeattle
Period16/06/2422/06/24

Funding

This research was funded by a Strategic Research Grant (Project No. 7005665) from City University of Hong Kong.

Research Keywords

  • Representation Learning
  • Video Object Tracking

Fingerprint

Dive into the research topics of 'Learning Tracking Representations from Single Point Annotations'. Together they form a unique fingerprint.

Cite this