TY - JOUR
T1 - Partial Domain Adaptation via Importance Sampling-Based Shift Correction
AU - Guo, Cheng-Jun
AU - Ren, Chuan-Xian
AU - Luo, You-Wei
AU - Xu, Xiao-Lin
AU - Yan, Hong
PY - 2025
Y1 - 2025
N2 - Partial domain adaptation (PDA) is a challenging task in real-world machine learning scenarios. It aims to transfer knowledge from a labeled source domain to a related unlabeled target domain, where the support set of the source label distribution subsumes the target one. Previous PDA works managed to correct the label distribution shift by weighting samples in the source domain. However, the simple reweighing technique cannot explore the latent structure and sufficiently use the labeled data, and then models are prone to over-fitting on the source domain. In this work, we propose a novel importance sampling-based shift correction (IS2C) method, where new labeled data are sampled from a built sampling domain, whose label distribution is supposed to be the same as the target domain, to characterize the latent structure and enhance the generalization ability of the model. We provide theoretical guarantees for IS2C by proving that the generalization error can be sufficiently dominated by IS2C. In particular, by implementing sampling with the mixture distribution, the extent of shift between source and sampling domains can be connected to generalization error, which provides an interpretable way to build IS2C. To improve knowledge transfer, an optimal transport-based independence criterion is proposed for conditional distribution alignment, where the computation of the criterion can be adjusted to reduce the complexity from O(n3) to O(n2) in realistic PDA scenarios. Extensive experiments on PDA benchmarks validate the theoretical results and demonstrate the effectiveness of our IS2C over existing methods. © 2025 IEEE.
AB - Partial domain adaptation (PDA) is a challenging task in real-world machine learning scenarios. It aims to transfer knowledge from a labeled source domain to a related unlabeled target domain, where the support set of the source label distribution subsumes the target one. Previous PDA works managed to correct the label distribution shift by weighting samples in the source domain. However, the simple reweighing technique cannot explore the latent structure and sufficiently use the labeled data, and then models are prone to over-fitting on the source domain. In this work, we propose a novel importance sampling-based shift correction (IS2C) method, where new labeled data are sampled from a built sampling domain, whose label distribution is supposed to be the same as the target domain, to characterize the latent structure and enhance the generalization ability of the model. We provide theoretical guarantees for IS2C by proving that the generalization error can be sufficiently dominated by IS2C. In particular, by implementing sampling with the mixture distribution, the extent of shift between source and sampling domains can be connected to generalization error, which provides an interpretable way to build IS2C. To improve knowledge transfer, an optimal transport-based independence criterion is proposed for conditional distribution alignment, where the computation of the criterion can be adjusted to reduce the complexity from O(n3) to O(n2) in realistic PDA scenarios. Extensive experiments on PDA benchmarks validate the theoretical results and demonstrate the effectiveness of our IS2C over existing methods. © 2025 IEEE.
KW - Partial domain adaptation
KW - importance sampling
KW - generalization error analysis
KW - label shift
KW - conditional shift
UR - https://www.webofscience.com/wos/woscc/full-record/WOS:001550831100002
UR - https://www.scopus.com/pages/publications/105012376460
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-105012376460&origin=recordpage
U2 - 10.1109/TIP.2025.3593115
DO - 10.1109/TIP.2025.3593115
M3 - RGC 21 - Publication in refereed journal
SN - 1057-7149
VL - 34
SP - 5009
EP - 5022
JO - IEEE Transactions on Image Processing
JF - IEEE Transactions on Image Processing
ER -