Abstract
Over the past decades, advances in artificial intelligence have introduced new opportunities and paradigms to the natural sciences, impacting areas such as data analysis, text mining, and hypothesis generation. This progress has led to the emergence of a new interdisciplinary field: AI for Science (AI4SCI). Within this domain, AI-driven drug discovery—including tasks such as DNA methylation detection and drug–target interaction prediction—has remained a prominent research focus. This thesis presents three studies, each addressing a key challenge in this area and contributing novel methodologies and insights.The first study proposes an interpretable deep learning approach for the prediction of four types of DNA methylation modifications. Although various computational methods have been developed for DNA methylation prediction, two main limitations persist: (1) existing models are typically confined to binary classification, which only assesses the presence or absence of modifications, thereby limiting comprehensive analysis of the interplay among different modification types; and (2) few studies provide adequate explanations of model decision-making, often relying solely on attention matrix visualization, which offers limited interpretability. To address these issues, this work frames DNA methylation modification prediction as a multi-class classification problem for the first time, introducing iResNetDM, a deep learning model that integrates residual networks (ResNet) with self-attention mechanisms. To the best of our knowledge, iResNetDM is the first model capable of distinguishing among four types of DNA methylation modifications. The model demonstrates strong performance across multiple modification types and effectively captures relationships between them. Moreover, we employ integrated gradients to enhance the interpretability of iResNetDM, successfully identifying multiple motifs and elucidating the model’s decision-making process. Notably, iResNetDM exhibits robustness and is able to uncover unique motifs for different methylation modifications. Comparative analysis of motifs across modification types reveals notable sequence similarities, suggesting potential regulatory roles in gene expression.
The second study presents a drug–target interaction (DTI) prediction framework that combines modified hierarchical molecular graphs (mHMG) with an improved convolutional block attention module (iCBAM). Accurate DTI prediction is fundamental to understanding drug mechanisms yet remains challenging, particularly for proteins or compounds absent from training datasets. Existing computational approaches often fail to capture critical molecular motifs or spatial protein information. To overcome these challenges, we introduce mHMG-DTI, a framework that leverages iCBAM for enhanced protein feature extraction and mHMGs for comprehensive molecular encoding. This hierarchical strategy captures detailed local structures and broader connectivity patterns, incorporating guiding knowledge to improve feature representation. Extensive experiments on four benchmark datasets, encompassing both classification and regression tasks, show that mHMG-DTI outperforms baseline models in most cases. These results highlight its potential to improve DTI prediction accuracy, thereby accelerating drug discovery and offering insights into drug resistance and side effects.
The third study proposes a drug repositioning framework that integrates multi-agent collaboration, retrieval-augmented generation (RAG), and Monte Carlo tree search (MCTS). Recent advances in large language models (LLMs) have shown promise for scientific applications such as drug repositioning; however, their effectiveness is limited when reasoning requires knowledge beyond pretraining data. Conventional methods—such as fine-tuning or standalone RAG—either incur high computational costs or underutilize structured scientific data. To address these limitations, we develop DrugMCTS, a novel framework that synergistically combines RAG, multi-agent collaboration, and MCTS for drug repositioning. DrugMCTS employs five specialized agents to retrieve and analyze molecular and protein information, enabling structured and iterative reasoning. Experimental results on the DrugBank and KIBA datasets demonstrate that DrugMCTS achieves substantially higher recall and robustness compared with both general-purpose LLMs and deep learning baselines. These findings underscore the importance of structured reasoning, agent-based collaboration, and feedback-driven search in advancing LLM applications for drug discovery.
Collectively, these studies advance the state of the art in AI for drug discovery by introducing novel, interpretable, and robust frameworks for key tasks, thereby contributing valuable methodologies and insights to the field.
| Date of Award | 12 Mar 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Yudai MATSUDA (Supervisor) & Linqi SONG (Co-supervisor) |
Cite this
- Standard