TY - JOUR
T1 - DDintensity
T2 - Addressing imbalanced drug-drug interaction risk levels using pre-trained deep learning model embeddings
AU - Xie, Weidun
AU - Chen, Xingjian
AU - Huang, Lei
AU - Zheng, Zetian
AU - Wang, Yuchen
AU - Zhang, Ruoxuan
AU - Zhang, Xiao
AU - Liu, Zhichao
AU - Peng, Chengbin
AU - Gullerova, Monika
AU - Wong, Ka-chun
PY - 2025/10
Y1 - 2025/10
N2 - Imbalanced datasets have been a persistent challenge in bioinformatics, particularly in the context of drug-drug interaction (DDI) risk level datasets. Such imbalance can lead to biased models that perform poorly on underrepresented classes. To address this issue, one strategy is to construct a balanced dataset, while another involves employing more advanced features and models. In this study, we introduce a novel approach called DDintensity, which leverages pre-trained deep learning models as embedding generators combined with LSTM-attention models to address the imbalance in DDI risk level datasets. We tested embeddings from various domains, including images, graphs, and textual corpus. Among these, embeddings generated by BioGPT achieved the highest performance, with an Area Under the Curve (AUC) of 0.97 and an Area Under the Precision-Recall curve (AUPR) of 0.92. Our model was trained on the DDinter and further validated using the MecDDI dataset. Additionally, case studies on chemotherapeutic drugs, DB00398 (Sorafenib) and DB01204 (Mitoxantrone) used in oncology, were conducted to demonstrate the specificity and effectiveness of the this methods. Our approach demonstrates high scalability across DDI modalities, as well as the discovery of novel interactions. In summary, we introduce DDIntensity as a solution for imbalanced datasets in bioinformatics with pre-trained deep-learning embeddings. © 2025 Published by Elsevier B.V.
AB - Imbalanced datasets have been a persistent challenge in bioinformatics, particularly in the context of drug-drug interaction (DDI) risk level datasets. Such imbalance can lead to biased models that perform poorly on underrepresented classes. To address this issue, one strategy is to construct a balanced dataset, while another involves employing more advanced features and models. In this study, we introduce a novel approach called DDintensity, which leverages pre-trained deep learning models as embedding generators combined with LSTM-attention models to address the imbalance in DDI risk level datasets. We tested embeddings from various domains, including images, graphs, and textual corpus. Among these, embeddings generated by BioGPT achieved the highest performance, with an Area Under the Curve (AUC) of 0.97 and an Area Under the Precision-Recall curve (AUPR) of 0.92. Our model was trained on the DDinter and further validated using the MecDDI dataset. Additionally, case studies on chemotherapeutic drugs, DB00398 (Sorafenib) and DB01204 (Mitoxantrone) used in oncology, were conducted to demonstrate the specificity and effectiveness of the this methods. Our approach demonstrates high scalability across DDI modalities, as well as the discovery of novel interactions. In summary, we introduce DDIntensity as a solution for imbalanced datasets in bioinformatics with pre-trained deep-learning embeddings. © 2025 Published by Elsevier B.V.
KW - Imbalanced datasets
KW - Deep learning
KW - Embeddings
KW - Drug-drug interaction predictions
UR - https://www.scopus.com/pages/publications/105009631481
U2 - 10.1016/j.artmed.2025.103202
DO - 10.1016/j.artmed.2025.103202
M3 - RGC 21 - Publication in refereed journal
SN - 0933-3657
VL - 168
JO - Artificial Intelligence in Medicine
JF - Artificial Intelligence in Medicine
M1 - 103202
ER -