Skip to main navigation Skip to search Skip to main content

Token Compensator: Altering Inference Cost of Vision Transformer Without Re-tuning

  • Shibo Jie
  • , Yehui Tang
  • , Jianyuan Guo
  • , Zhi-Hong Deng*
  • , Kai Han*
  • , Yunhe Wang*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Token compression expedites the training and inference of Vision Transformers (ViTs) by reducing the number of the redundant tokens, e.g., pruning inattentive tokens or merging similar tokens. However, when applied to downstream tasks, these approaches suffer from significant performance drop when the compression degrees are mismatched between training and inference stages, which limits the application of token compression on off-the-shelf trained models. In this paper, we propose a model arithmetic framework to decouple the compression degrees between the two stages. In advance, we additionally perform a fast parameter-efficient self-distillation stage on the pre-trained models to obtain a small plugin, called Token Compensator (ToCom), which describes the gap between models across different compression degrees. During inference, ToCom can be directly inserted into any downstream off-the-shelf models with any mismatched training and inference compression degrees to acquire universal performance improvements without further training. Experiments on over 20 downstream tasks demonstrate the effectiveness of our framework. On CIFAR100, fine-grained visual classification, and VTAB-1k benchmark, ToCom can yield up to a maximum improvement of 2.3%, 1.5%, and 2.0% in the average performance of DeiT-B, respectively. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings, Part XVI
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer, Cham
Pages76-94
ISBN (Electronic)9783031726408
ISBN (Print)9783031726392
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event18th European Conference on Computer Vision (ECCV 2024) - MiCo Milano, Milan, Italy
Duration: 29 Sept 20244 Oct 2024
https://eccv.ecva.net/

Publication series

NameLecture Notes in Computer Science
Volume15074
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision (ECCV 2024)
Abbreviated titleECCV2024
PlaceItaly
CityMilan
Period29/09/244/10/24
Internet address

Fingerprint

Dive into the research topics of 'Token Compensator: Altering Inference Cost of Vision Transformer Without Re-tuning'. Together they form a unique fingerprint.

Cite this