Skip to main navigation Skip to search Skip to main content

Rethinking Pruning for Accelerating Deep Inference at the Edge

  • Dawei Gao
  • , Xiaoxi He
  • , Zimu Zhou*
  • , Yongxin Tong
  • , Ke Xu
  • , Lothar Thiele
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

There is a growing trend to deploy deep neural networks at the edge for high-accuracy, real-time data mining and user interaction. Applications such as speech recognition and language understanding often apply a deep neural network to encode an input sequence and then use a decoder to generate the output sequence. A promising technique to accelerate these applications on resource-constrained devices is network pruning, which compresses the size of the deep neural network without severe drop in inference accuracy. However, we observe that although existing network pruning algorithms prove effective to speed up the prior deep neural network, they lead to dramatic slowdown of the subsequent decoding and may not always reduce the overall latency of the entire application. To rectify such drawbacks, we propose entropy-based pruning, a new regularizer that can be seamlessly integrated into existing network pruning algorithms. Our key theoretical insight is that reducing the information entropy of the deep neural network outputs decreases the upper bound of the subsequent decoding search space. We validate our solution with two state-of-the-art network pruning algorithms on two model architectures. Experimental results show that compared with existing network pruning algorithms, our entropy-based pruning method notably suppresses and even eliminates the increase of decoding time, and achieves shorter overall latency with only negligible extra accuracy loss in the applications. © 2020 ACM.
Original languageEnglish
Title of host publicationKDD 2020 - Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
PublisherAssociation for Computing Machinery
Pages155-164
ISBN (Print)9781450379984
DOIs
Publication statusPublished - Aug 2020
Externally publishedYes
Event26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2020) - Virtual, CA, United States
Duration: 23 Aug 202027 Aug 2020
https://www.kdd.org/kdd2020/

Publication series

NameProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

Conference

Conference26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2020)
PlaceUnited States
CityCA
Period23/08/2027/08/20
Internet address

Funding

We thank the anonymous reviewers for their valuable suggestions and comments. Dawei Gao, Yongxin Tong and Ke Xu’s work was partially supported by the National Key Research and Development Program of China under Grant No. 2018AAA0101100, the National Science Foundation of China (NSFC) under Grant No. 61822201, U1811463 and 71531001, and the Beijing Municipal Science and Technology Project under Grant Z191100002519012. Xiaoxi He and Lothar Thiele’s work was supported in part by the Swiss National Science Foundation in the context of the NCCR Automation. Zimu Zhou’s research was supported in part by the Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 1 grant.

Research Keywords

  • automatic speech recognition
  • deep learning
  • name entity recognition
  • network pruning
  • sequence labelling

Fingerprint

Dive into the research topics of 'Rethinking Pruning for Accelerating Deep Inference at the Edge'. Together they form a unique fingerprint.

Cite this