Skip to main navigation Skip to search Skip to main content

Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

16 Downloads (CityUHK Scholars)

Abstract

Large language models (LLMs) have made exciting achievements across various domains, yet their deployment on resource-constrained personal devices remains hindered by the prohibitive computational and memory demands of task-specific fine-tuning. While quantization offers a pathway to efficiency, existing methods struggle to balance performance and overhead, either incurring high computational/memory costs or failing to address activation outliers, a critical bottleneck in quantized fine-tuning. To address these challenges, we propose the Outlier Spatial Stability Hypothesis (OSSH): During fine-tuning, certain activation outlier channels retain stable spatial positions across training iterations. Building on OSSH, we propose Quaff, a Quantized parameter-efficient fine-tuning framework for LLMs, optimizing low-precision activation representations through targeted momentum scaling. Quaff dynamically suppresses outliers exclusively in invariant channels using lightweight operations, eliminating full-precision weight storage and global rescaling while reducing quantization errors. Extensive experiments across ten benchmarks validate OSSH and demonstrate Quaff's efficacy. Specifically, on the GPQA reasoning benchmark, Quaff achieves a 1.73× latency reduction and 30% memory savings over full-precision fine-tuning while improving accuracy by 0.6% on the Phi-3 model, reconciling the triple trade-off between efficiency, performance, and deployability. By enabling consumer-grade GPU fine-tuning (e.g., RTX 2080 Super) without sacrificing model utility, Quaff democratizes personalized LLM deployment. The code is available at https://github.com/Little0o0/Quaff.git. © 2025 Association for Computational Linguistics.  
Original languageEnglish
Title of host publicationProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics
EditorsWanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
PublisherAssociation for Computational Linguistics
Pages6481-6496
Number of pages16
Volume1
ISBN (Print)979-8-89176-251-0
DOIs
Publication statusPublished - 2025
Event63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025) - Austria Center Vienna, Vienna, Austria
Duration: 27 Jul 20251 Aug 2025
https://2025.aclweb.org/
https://aclanthology.org/2025.acl-long/
https://aclanthology.org/volumes/2025.findings-acl/

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
Volume1
ISSN (Print)0736-587X

Conference

Conference63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)
PlaceAustria
CityVienna
Period27/07/251/08/25
Internet address

Funding

This paper is partially supported by Hong Kong Research Grants Council (RGC) grant #11203523.

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis'. Together they form a unique fingerprint.
  • GRF: Resource-Constrained Federated Learning

    WU, D. (Principal Investigator / Project Coordinator)

    1/09/23 → …

    Project: Research

Cite this