High-dimensional data analysis has been attracting considerable attention, since most real applications
include high-dimensional data, e.g., gene expression data and global climate patterns, etc. Researchers
working in these areas usually require confronting with the issue of how to represent high-dimensional
data, such as extracting informative features holding all useful information. This problem can benefit
from dimensionality reduction (DR) or manifold learning, which converts high-dimensional data into a
low-dimensional feature subspace in which certain important characteristics of data are preserved. The
ever-increasing multimedia (such as images) data in daily communication and emerging applications
poses new challenge of effective high-dimensional data understanding, so data interpretation remains a
difficult issue. In this thesis, we mainly focus on the study of data representation. More specifically, we
propose four novel discriminant dimensionality reduction techniques, along with illustrating the results
on benchmark problems, for feature extraction, visualization and classification.
First, we incorporate the pairwise constraints into the well-known Isomap framework and propose an
orthogonal Marginal Isomap (M-Isomap) for discriminant nonlinear DR. In this setting, the pairwise
constraints are applied to specify the types of neighborhoods. M-Isomap computes the shortest path
distances over constrained neighborhood graphs and guides nonlinear DR by separating the inter-class
neighbors. Thus, large margins between both inter- and intra-class clusters are delivered and enhanced
compactness of intra-cluster points is achieved at the same time. The validity of M-Isomap is examined
by extensive multivariate visualization and clustering evaluation on synthetic and benchmark databases.
The results are compared with other related state-of-the-art manifold learning algorithms.
Then, we propose a Constrained Large Margin Local Projection (CMLP) for multimodal linear DR
and data classification. We elaborate the CMLP criterion from a constrained marginal perspective. Four
effective solution schemes are presented. CMLP is originated from Locality Preserving Projections
(LPP), but it offers certain advantages over LPP. In order to reflect the proximity relations of inter- and
intra-class neighboring pairs, the localized pairwise constraints are employed to specify the types of
neighbors. By using CMLP, margins between inter- and intra-class clusters are significantly enlarged. So, multimodal structures can be effectively preserved. Extensive data visualization and classification
on benchmark datasets are conducted to verify the efficiency of our techniques.
Considering that supervised DR for classification may suffer from the overfitting if labeled number
is few, we further propose two sub-manifold projection based semi-supervised DR algorithms termed
Marginal Semi-Supervised Sub-Manifold Projections (MS3MP) and orthogonal MS3MP (OMS3MP),
which are semi-supervised extensions of CMLP with different graph weight construction. Based on the
constructed manifold scatters, our techniques can preserve local information of all samples (including
labeled, constrained and unlabeled) and discriminant structures embedded in the localized PC. The
sub-manifolds of different classes can also be separated. Considering that in all PC guided algorithms,
determining the informative constraint subset is challenging and random subsets significantly affect the
performance of the algorithms, we also introduce an effective technique as a guide in selecting the
informative constraints for DR with consistent constraints. The connections between this present work
and other works are also elaborated. Benchmark simulations demonstrate the validity of our algorithms.
Finally, we discuss sparse representation for DR. By incorporating the group sparse representation
into the well-known Canonical Correlation Analysis (CCA) framework, we technically propose a new
DR algorithm named Group Sparse Canonical Correlation Analysis (GSCCA) for discriminant feature
extraction and classification. GSCCA uses two sets of variables and aims to preserve the group sparse
characteristics of data within each set in addition to maximize the global inter-set covariance. With GS
weights calculated before DR and feature extraction, the locality, sparsity and discriminant information
of data can be adaptively determined. The GS weights are calculated from a NP-hard group-sparsity
promoting problem that considers all highly correlated data within a group. By setting one of the two
variable sets as the class label matrix, GSCCA is extended to multi-class case. The projection matrix of
GSCCA can be analytically obtained. Extensive simulations are conducted to examine GSCCA.
| Date of Award | 2 Oct 2013 |
|---|
| Original language | English |
|---|
| Awarding Institution | - City University of Hong Kong
|
|---|
| Supervisor | Wai Shing Tommy CHOW (Supervisor) & Chi Wai Derek PAO (Supervisor) |
|---|
- Image processing
- Dimension reduction (Statistics)
The study of discriminant dimensionality reduction for feature extraction, visualization and classification
ZHANG, Z. (Author). 2 Oct 2013
Student thesis: Doctoral Thesis