Abstract
With the rapid development of big data, tensors, which refer to multi-dimensional arrays, are widely used in various fields, including neuroscience, image analysis, climatology, and spatiotemporal forecasting. Tensor regression, which is a powerful technique for solving high dimensional problems, has received significant attention in recent years. To address the structural complexity of tensors with high order, tensor train decomposition offers an efficient method for imposing low-rank constraints on the coefficient tensor. Meanwhile, the single-machine system is difficult to store and process large datasets; thus, distributed frameworks have been developed to enable efficient analysis across multiple machines. In this thesis, we study tensor train regression with tensor predictors and scalar responses, utilizing convex regularization, and extend it to a distributed framework (communication-efficient surrogate likelihood framework). Then we consider the decentralized distributed quantile regression over networks. The main contents and conclusions of this thesis are listed as follows.In Chapter 2, we study tensor regression with tensor predictors and scalar responses based on tensor train (TT) decomposition using nuclear norm penalties. The statistical rates of the estimators are established in the high-dimensional setting, for both mean regression and quantile regression. Numerical experiments and empirical applications are presented to demonstrate their finite-sample performance. This work provides theoretical guarantees for a stable and efficient alternative for high-dimensional tensor regression based on the TT format, which scales well for high-order tensors.
In Chapter 3, we consider the distributed learning problem for regularized tensor train regression with a low TT-rank structure. We extend the tensor train regression model to the communication-efficient surrogate likelihood framework, while the distributed framework imposes the surrogate loss function based on Taylor’s expression. We establish the upper bound of the distributed estimator and show that our distributed estimators reach the same convergence rate as the global estimator under a suitable initial value and certain assumptions. Our method extends the theoretical properties of distributed tensor regression, and some numerical experiments are conducted to illustrate the finite sample performance of the proposed method.
In Chapter 4, we present a distributed quantile regression framework that integrates a fused lasso penalty over a K-Nearest-Neighbors (K-NN) graph. The K-NN graph is constructed using adaptive metrics, such as spatial distance or model similarity. With the subgraph, our proposed model could reduce computational complexity while preserving estimation accuracy. We establish theoretical guarantees for the proposed estimator, including the asymptotic normality and consistency in variable selection. To solve the optimization problem efficiently, we develop a distributed Alternating Direction Method of Multipliers (ADMM) algorithm. Numerical experiments on both synthetic and real-world datasets demonstrate the superior performance of our method in terms of estimation accuracy and cluster identification compared to existing graph-based approaches.
| Date of Award | 19 May 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Heng LIAN (Supervisor) |
Cite this
- Standard