Skip to main navigation Skip to search Skip to main content

Extreme Image Compression Using Fine-tuned VQGANs

  • Qi Mao*
  • , Tinghan Yang
  • , Yinuo Zhang
  • , Zijian Wang
  • , Meng Wang
  • , Shiqi Wang
  • , Libiao Jin
  • , Siwei Ma
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Recent advances in generative compression methods have demonstrated remarkable progress in enhancing the perceptual quality of compressed data, especially in scenarios with low bitrates. However, their efficacy and applicability to achieve extreme compression ratios (< 0.05 bpp) remain constrained. In this work, we propose a simple yet effective coding framework by introducing vector quantization (VQ)-based generative models into the image compression domain. The main insight is that the codebook learned by the VQGAN model yields a strong expressive capacity, facilitating efficient compression of continuous information in the latent space while maintaining reconstruction quality. Specifically, an image can be represented as VQ-indices by finding the nearest codeword, which can be encoded using lossless compression methods into bitstreams. We propose clustering a pre-trained large-scale codebook into smaller codebooks through the K-means algorithm, yielding variable bitrates and different levels of reconstruction quality within the coding framework. Furthermore, we introduce a transformer to predict lost indices and restore images in unstable environments. Extensive qualitative and quantitative experiments on various benchmark datasets demonstrate that the proposed framework outperforms state-of-the-art codecs in terms of perceptual quality-oriented metrics and human perception at extremely low bitrates (≤ 0.04 bpp). Remarkably, even with the loss of up to 20% of indices, the images can be effectively restored with minimal perceptual loss. © 2024 IEEE.
Original languageEnglish
Title of host publicationProceedings - DCC 2024: 2024 Data Compression Conference
EditorsAli Bilgin, James E. Fowler, Joan Serra-Sagrista, Yan Ye, James A. Storer
PublisherIEEE
Pages203-212
ISBN (Electronic)9798350385878
ISBN (Print)979-8-3503-8588-5
DOIs
Publication statusPublished - 2024
Event2024 Data Compression Conference (DCC 2024) - Snowbird, United States
Duration: 19 Mar 202422 Mar 2024

Publication series

NameData Compression Conference Proceedings
ISSN (Print)1068-0314
ISSN (Electronic)2375-0359

Conference

Conference2024 Data Compression Conference (DCC 2024)
PlaceUnited States
CitySnowbird
Period19/03/2422/03/24

Bibliographical note

Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).

Funding

This work was supported in part by the National Natural Science Foundation of China under Grants 62201526, 62025101, and 62022002; the Fundamental Research Funds for the Central Universities (CUC23GZ007); and the Public Computing Cloud at CUC, all of which are gratefully acknowledged.

Fingerprint

Dive into the research topics of 'Extreme Image Compression Using Fine-tuned VQGANs'. Together they form a unique fingerprint.

Cite this