Abstract
With the rapid growth of large language models (LLMs), there is increasing demand for memory and computation in LLMs. Recent efforts on post-training pruning of LLMs aim to reduce the model size and computation requirements, yet the performance is still sub-optimal. In this paper, we present a plug-and-play solution for post-training pruning of LLMs. The proposed solution has two innovative components: 1) Relative Importance and Activations (RIA), a new pruning metric that jointly considers the weight and activations efficiently on LLMs, and 2) Channel Permutation, a new approach to maximally preserves important weights under N:M sparsity. The two proposed components can be readily combined to further enhance the N:M semi-structured pruning of LLMs. Our empirical experiments show that RIA alone can already surpass all existing post-training pruning methods on prevalent LLMs, e.g., LLaMA ranging from 7B to 65B. Furthermore, N:M semi-structured pruning with channel permutation can even outperform the original LLaMA2-70B on zero-shot tasks, together with practical speed-up on specific hardware. Our code is available at: https://github.com/biomedical-cybernetics/Relative-importance-and-activation-pruning
| Original language | English |
|---|---|
| Title of host publication | The Twelfth International Conference on Learning Representations |
| Subtitle of host publication | ICLR 2024 |
| Number of pages | 19 |
| Publication status | Published - May 2024 |
| Externally published | Yes |
| Event | 12th International Conference on Learning Representations (ICLR 2024) - Messe Wien Exhibition and Congress Center, Vienna, Austria Duration: 7 May 2024 → 11 May 2024 https://iclr.cc/Conferences/2024 https://openreview.net/group?id=ICLR.cc/2024/Conference |
Conference
| Conference | 12th International Conference on Learning Representations (ICLR 2024) |
|---|---|
| Place | Austria |
| City | Vienna |
| Period | 7/05/24 → 11/05/24 |
| Internet address |
Fingerprint
Dive into the research topics of 'Plug-and-Play: An Efficient Post-training Pruning Method for Large Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver