Abstract
Identifying potential compounds for target proteins is crucial in drug discovery. Current compound–protein interaction prediction models rely on complex features to enhance capabilities, but this often incurs substantial computational burdens. Indeed, this challenge arises from the limited understanding of data imbalance between proteins and compounds, leading to insufficient optimization of protein encoders. To address this issue, a sequence-based predictor named FilmCPI is introduced, which leverages data imbalance by learning proteins with their numerous corresponding compounds. This approach enables the characteristics of each protein to be effectively represented through itself and its relationship with corresponding compounds. Without increasing parameters, FilmCPI consistently outperforms baseline models across diverse datasets and split strategies, and its generalization to unseen proteins becomes more pronounced as the datasets expand. Notably, FilmCPI can be effectively transferred to unseen membrane protein families with sequence-based data from other families. The optimization dynamics is further analyzed and it is discovered that the effectiveness of FilmCPI is attributed to different optimization speeds for diverse encoders. Overall, this work aims to provide a theoretical perspective for designing efficient models. © 2025 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH.
| Original language | English |
|---|---|
| Article number | 2400985 |
| Number of pages | 9 |
| Journal | Advanced Intelligent Systems |
| Volume | 7 |
| Issue number | 12 |
| Online published | 7 Jul 2025 |
| DOIs | |
| Publication status | Published - Dec 2025 |
Funding
This work is supported by a start-up grant (grant no. 9610591) for New Faculty and an internal grant (grant no. 7006055) from the City University of Hong Kong to C.C.A.F. and the general grant (project no. JCYJ20230807115001004) from the Science, Technology, and Innovation Commission of Shenzhen Municipality to C.C.A.F. (The Shenzhen Research Institute, City University of Hong Kong).
Research Keywords
- compound–protein interaction
- data imbalance
- drug discovery
- feature-wise linear modulation
- neural network
Publisher's Copyright Statement
- This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/
Fingerprint
Dive into the research topics of 'Utilizing Data Imbalance to Enhance Compound–Protein Interaction Prediction Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver