Abstract
Cancer is a large group of genetic diseases, featuring both inter- and intra-tumor heterogeneity. Recent technologies in genomic profiling have provided great opportunities to investigate the molecular changes during tumor development and progression. At the same time, numerous pattern classification techniques have been developed to process and interpret high-throughput molecular data. Cancer subtyping is an emerging subject area with the final goal of dissecting tumor heterogeneity. Previous studies in different cancer types, such as breast and colorectal cancer, have allowed the separation of patients into molecularly and clinically homogeneous subtypes. However, for other cancer types, such as nasopharyngeal carcinoma (NPC) and pancreatic ductal adenocarcinoma (PDAC), relatively few or no subtyping studies have been conducted. Although highly heterogeneous characteristics have been observed in both NPC and PDAC, patients diagnosed with NPC or PDAC are still treated as having a single disease. This thesis focuses on dissecting the heterogeneity of NPC and PDAC, with the goals of building classifiers which can be used to stratify patients into potentially homogeneous subtypes, and finding biomarkers associated with tumorigenesis. In order to pave the way toward precision medicine, new algorithms were developed to integrate analyses of multi-omics cancer data.NPC is very invasive and highly metastatic with diverse molecular characteristics and clinical outcomes. The aims of the first part of the thesis include dissect the heterogeneity, and construct a prognostic model for prediction of the likelihood of distant metastases of NPC. The thesis proposed, for the first time, that NPC can be stratified into three subtypes by consensus clustering. Using a panel of 4 microRNAs (miRNAs), a prognostic model was established that can robustly stratify NPC patients into high- and low- risk groups of developing distant metastases. To dissect the intra-tumor heterogeneity of PDAC, a retrospective meta-analysis on whole transcriptome data from more than 1,200 PDAC patients were performed. Unsupervised non-negative matrix factorization (NMF), a dimensionality reduction and factorization-based biclustering algorithm, was used to do cluster analysis. This is the largest cohort of PDAC gene expression profiles investigated so far, which greatly increased the statistical power of the analysis and provided more robust subtyping results. To gain more insights into the regulatory mechanism, and to identify master regulators of NPC, the first time-series analysis in NPC was carried out. The expression data were arranged into tensor forms, and a newly developed tensor decomposition algorithm called Sparse Decomposition of Arrays (SDA) was used to extract latent factors in the data. Regulatory networks were built for tumor and stroma, respectively. Moreover, two master regulators, which may drive the process of NPC tumorigenesis, were identified and validated by quantitative polymerase chain reaction (qPCR) experiments.
In recent years, more multi-omics datasets have emerged and they are needed to take into account toward a more systematic subtyping of individual cancers. How to take full advantage of the availability of these data, to do a systematic subtyping of cancer is the major focus of the later part of the thesis. Two algorithms were developed to meet this need. The first algorithm for cancer subtyping is called Joint Tensor and Matrices Decomposition (JTMD). JTMD was applied to integrate analyses of multi-omics data from The Cancer Genome Atlas (TCGA) with associated side information to grouping, and predicting clinical outcomes of cancer patients. The second algorithm called Multi-omics Network Integration and Consensus Clustering (MNICC) was built to classify cancers into homogeneous subtypes. Comparison to state-of-the-art approaches, such as Similarity Network Fusion (SNF) and Affinity Network Fusion (ANF), shows that MNICC can achieve better subtyping results in terms of obtaining more biologically meaningful and clinically distinct subtypes of cancer.
In summary, biomarkers were identified, and pattern classification models were built for NPC and PDAC, respectively, which shed new light on tumorigenesis and may have great clinical applications in personalized treatment for cancer patients. Besides, novel computational methods for cancer subtyping were developed. These techniques can be explored in the future to build software platforms for biomedical research and clinical applications.
| Date of Award | 16 May 2019 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Hong YAN (Supervisor) |
Cite this
- Standard