Skip to main navigation Skip to search Skip to main content

Interpretable Machine Learning Applications from Drug Synthesis to Medical Treatment

Student thesis: Doctoral Thesis

Abstract

The concept of health is essential in the assessment of an individual’s state. Indeed, it is associated with the entire cycle from the prevention of disease, disease diagnosis, drug discovery, personalized treatment recommendation, to other therapies, shaping the maintenance of quality life. Cancer, especially, has received extensive attention for its challenging mortality rate. In particular, metastatic cancer involves the spread of the primary tumor to other body organs. Without undermining the fact that early detection of cancer is essential, the timely detection of metastasis, the identification of biomarkers, the development of suitable drugs, and the recommendation of treatment choice are univocally valuable.

In the first section, among standard drug discovery procedures, the synthesis of bioactive compounds from plants is examined under the lens of machine learning methods with metaheuristic algorithms. Bioactive compounds in plants can be synthesized using N-arylation methods such as the Buchwald-Hartwig reaction; it is essential in cancer drug discovery for their pharmacological effects. Specifically, important descriptors are necessary for the estimation of those reaction yields. We explore ten metaheuristic algorithms for descriptor selection and propose a voting ensemble. The algorithms were evaluated based on computational time and the number of selected descriptors. The analyses reflect that robust performance is obtained with additional descriptors. In addition, the essential descriptor was deduced based on the frequency of its occurrence within the 50 extracted data subsets. The voting algorithm, which is an ensemble method, was modeled and compared with other algorithms based on the Root Mean Square Error (RMSE) and R-square (𝑅2) evaluation metric. Important descriptors for yield estimation during chemical synthesis are noted for interpretability.

In the second section, we take a critical look at one of the key concerns in metastatic cancer treatment using machine learning (ML). The main objective here is to determine treatment discontinuation, with a particular focus on metastatic castration-resistant prostate cancer. According to the World Health Organization (WHO), prostate cancer is the second leading cause of death in men worldwide. Therefore, its prevention and treatment demand serious attention. Thus, we propose a Particle Swarm Optimized Gaussian Process Classifier (PSO-GPC) for the determination of treatment discontinuation. Based on three cohorts of prostate cancer patients, several classifiers were compared with PSO-GPC for their performance in determining treatment discontinuation. Given the data skewness and class imbalance, the models were evaluated based on both the area under receiver operating characteristics curve (AUROC) and area under precision recall curve (AUPRC). The proposed model, PSO-GPC, performs better than the state-of-the-art methods as well as the traditional machine learning methods. In addition, statistical analysis was carried out for ranking the methods and an independent cohort data was analyzed with PSO-GPC, demonstrating its unbiased performance. Furthermore, the factors that contribute to cancer and treatment discontinuation were interpreted and expatiated. A proper determination of treatment discontinuation in metastatic castration-resistant prostate cancer patients will reduce the mortality rate in cancer patients.

Machine learning and other computational tools can sometimes be regarded as “black-box” techniques, partly due to the difficulties in associating their internal workings with model outputs. Therefore, in the last section, we investigate the adoption of causal inference and SHAP insight in the context of tree-based machine learning algorithms for treatment continuation prediction. SHAP provides insights into the attributes which the algorithms put into consideration for prediction. It gives us an approach illuminating the opacity of “black-box” with certain intricacies of the decision-making process. In addition, given the genetic heterogeneity among individuals, population-based generalization might undermine the accuracy and put the application of our predictions at risk. On this note, SHAP can be used to provide local explainability for each patient. Furthermore, causal inference was used to determine the effect of smoking on treatment discontinuation, as well as the attribute which triggers adverse events. Therefore, the integration of causal inference and SHAP insights with machine learning algorithms can provide a reasonable and interpretable approach for treatment in broader contexts.
Date of Award22 Aug 2023
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorKa Chun WONG (Supervisor)

Cite this

'