Skip to main navigation Skip to search Skip to main content

Approximation Theory of Deep Neural Networks and Their Applications with Tight Framelets

Student thesis: Doctoral Thesis

Abstract

Deep neural networks (DNNs) have achieved remarkable success in a wide range of practical applications, attracting significant attention from researchers across various fields of science and technology. A key aspect of their appeal lies in their exceptional expressive power, which has a background in the approximation theory. Another important aspect is to enhance their performance with explainable mathematical techniques, such as wavelets/framelets, which have shown great success in signal processing. This thesis aims to address these challenges and provide solutions to the problems above.

Activation functions define how neurons in DNNs process incoming signals for them. They are essential for learning non-linear transformations and performing diverse computations among successive neuron layers. In the last few years, researchers have investigated the approximation ability of DNNs to explain their power and success. In this part, we explore the approximation ability of DNNs using a different activation function called SignReLU. Our theoretical results demonstrate that SignReLU networks outperform rational and ReLU networks in terms of approximation performance. Numerical experiments compare SignReLU with existing activations such as ReLU, LeakyReLU, and ELU, illustrating the competitive practical performance of SignReLU.

In the second part, we establish some analysis for linear feature extraction by deep convolutional neural networks (DCNNs), which demonstrates the power of deep learning over traditional linear transformations, like Fourier, wavelets, and redundant dictionary coding methods. Moreover, we explain how linear feature extraction can be conducted efficiently with multi-channel DCNNs. It can be applied to lower the essential dimension for approximating a high-dimensional function. Rates of function approximation by such deep networks implemented with channels and followed by fully connected layers are investigated as well. Analysis for factorizing linear features into multi-resolution convolutions is essential in our work. Nevertheless, a dedicated vectorization of matrices is constructed, which bridges 1-D DCNN and 2-D DCNN and allows us to have corresponding 2D analysis.

Finally, from the perspective of signal processing, we combine DCNNs with Haar-type framelets to improve the performance of deep neural networks. The nature of signals on the sphere and graphs differs from those defined in the Euclidean spaces and causes difficulties in designing wavelets/framelets and deep learning models. In this part, we develop a general theoretical framework for constructing Haar-type tight framelets on any compact set with a hierarchical partition. In particular, we build an area-regular hierarchical partition on the 2-sphere and establish its corresponding spherical Haar tight framelets with directionality. We also extend the proposed method to framelets on graphs and provide a theoretical analysis of their properties, such as permutation equivariance, efficiency, and sparsity, which are essential for practical purposes. In the end, we conclude by evaluating and illustrating the effectiveness of our area-regular spherical Haar tight framelets in several denoising experiments. Furthermore, we propose a convolutional neural network (CNN) model for spherical signal denoising that employs fast framelet decomposition and reconstruction algorithms. Experimental results show that our proposed CNN model outperforms threshold methods and processes strong generalization and robustness properties.
Date of Award19 Aug 2024
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorHan FENG (Supervisor) & Xiaosheng ZHUANG (Co-supervisor)

Cite this

'