Skip to main navigation Skip to search Skip to main content

Utilizing Data Imbalance to Enhance Compound–Protein Interaction Prediction Models

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

11 Downloads (CityUHK Scholars)

Abstract

Identifying potential compounds for target proteins is crucial in drug discovery. Current compound–protein interaction prediction models rely on complex features to enhance capabilities, but this often incurs substantial computational burdens. Indeed, this challenge arises from the limited understanding of data imbalance between proteins and compounds, leading to insufficient optimization of protein encoders. To address this issue, a sequence-based predictor named FilmCPI is introduced, which leverages data imbalance by learning proteins with their numerous corresponding compounds. This approach enables the characteristics of each protein to be effectively represented through itself and its relationship with corresponding compounds. Without increasing parameters, FilmCPI consistently outperforms baseline models across diverse datasets and split strategies, and its generalization to unseen proteins becomes more pronounced as the datasets expand. Notably, FilmCPI can be effectively transferred to unseen membrane protein families with sequence-based data from other families. The optimization dynamics is further analyzed and it is discovered that the effectiveness of FilmCPI is attributed to different optimization speeds for diverse encoders. Overall, this work aims to provide a theoretical perspective for designing efficient models. © 2025 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH.
Original languageEnglish
Article number2400985
Number of pages9
JournalAdvanced Intelligent Systems
Volume7
Issue number12
Online published7 Jul 2025
DOIs
Publication statusPublished - Dec 2025

Funding

This work is supported by a start-up grant (grant no. 9610591) for New Faculty and an internal grant (grant no. 7006055) from the City University of Hong Kong to C.C.A.F. and the general grant (project no. JCYJ20230807115001004) from the Science, Technology, and Innovation Commission of Shenzhen Municipality to C.C.A.F. (The Shenzhen Research Institute, City University of Hong Kong).

Research Keywords

  • compound–protein interaction
  • data imbalance
  • drug discovery
  • feature-wise linear modulation
  • neural network

Publisher's Copyright Statement

  • This full text is made available under CC-BY 4.0. https://creativecommons.org/licenses/by/4.0/

Fingerprint

Dive into the research topics of 'Utilizing Data Imbalance to Enhance Compound–Protein Interaction Prediction Models'. Together they form a unique fingerprint.

Cite this