Skip to main navigation Skip to search Skip to main content

Plug-and-Play: An Efficient Post-training Pruning Method for Large Language Models

  • Yingtao Zhang*
  • , Haoli Bai
  • , Haokun Lin
  • , Jialin Zhao
  • , Lu Hou
  • , Carlo Vittorio Cannistraci
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

With the rapid growth of large language models (LLMs), there is increasing demand for memory and computation in LLMs. Recent efforts on post-training pruning of LLMs aim to reduce the model size and computation requirements, yet the performance is still sub-optimal. In this paper, we present a plug-and-play solution for post-training pruning of LLMs. The proposed solution has two innovative components: 1) Relative Importance and Activations (RIA), a new pruning metric that jointly considers the weight and activations efficiently on LLMs, and 2) Channel Permutation, a new approach to maximally preserves important weights under N:M sparsity. The two proposed components can be readily combined to further enhance the N:M semi-structured pruning of LLMs. Our empirical experiments show that RIA alone can already surpass all existing post-training pruning methods on prevalent LLMs, e.g., LLaMA ranging from 7B to 65B. Furthermore, N:M semi-structured pruning with channel permutation can even outperform the original LLaMA2-70B on zero-shot tasks, together with practical speed-up on specific hardware. Our code is available at: https://github.com/biomedical-cybernetics/Relative-importance-and-activation-pruning
Original languageEnglish
Title of host publicationThe Twelfth International Conference on Learning Representations
Subtitle of host publicationICLR 2024
Number of pages19
Publication statusPublished - May 2024
Externally publishedYes
Event12th International Conference on Learning Representations (ICLR 2024) - Messe Wien Exhibition and Congress Center, Vienna, Austria
Duration: 7 May 202411 May 2024
https://iclr.cc/Conferences/2024
https://openreview.net/group?id=ICLR.cc/2024/Conference

Conference

Conference12th International Conference on Learning Representations (ICLR 2024)
PlaceAustria
CityVienna
Period7/05/2411/05/24
Internet address

Fingerprint

Dive into the research topics of 'Plug-and-Play: An Efficient Post-training Pruning Method for Large Language Models'. Together they form a unique fingerprint.

Cite this