Skip to main navigation Skip to search Skip to main content

Towards Secure, Efficient, and Privacy-preserving Federated Learning

Student thesis: Doctoral Thesis

Abstract

Federated Learning (FL) has emerged as a powerful paradigm for collaborative model training without requiring direct data sharing, enabling both privacy preservation and broad applicability across domains. However, the widespread deployment of FL remains challenged by three critical issues: security, efficiency, and the right to privacy in data lifecycle management. This dissertation addresses these challenges through three complementary contributions.

First, we investigate the vulnerability of Byzantine-robust FL systems to novel poisoning threats. While prior work assumes that malicious updates are distinguishable from benign ones, we introduce BadSampler, a clean-label poisoning attack that harnesses catastrophic forgetting to degrade the generalization ability of Byzantine-robust FL. We formally establish the connection between catastrophic forgetting and generalization error, design efficient adversarial sampling strategies, and provide theoretical guarantees alongside extensive empirical validation, demonstrating the effectiveness of the proposed attack even under strong defenses.

Second, we improve the computational efficiency of Vertical Federated Learning (VFL), where heterogeneous participants collaborate under split-learning settings. To overcome the limitations of synchronous dependencies and heterogeneity-induced inefficiencies, we propose PubSub-VFL, a novel Publisher/Subscriber-based paradigm that integrates hierarchical asynchrony with data-parallel parameter server architectures. Our approach supports privacy-preserving optimization, achieves stable convergence, and significantly reduces training latency. Experimental results on multiple real-world datasets show that PubSub-VFL accelerates training by up to 2 ~ 7 x while preserving accuracy and achieving over 90% resource utilization.

Finally, we tackle the emerging challenge of machine unlearning in FL, motivated by the right to be forgotten. Unlike centralized unlearning methods that require full data access, we design an efficient rapid retraining framework that enables participants to collaboratively erase data contributions without exposing raw data. Our formal analysis establishes convergence and computational efficiency guarantees, while experiments on diverse datasets confirm that the proposed method achieves effective unlearning with minimal overhead and negligible utility loss.

Together, these contributions advance the state of the art in secure, efficient, and privacy-preserving federated learning. By systematically addressing poisoning resilience, computational efficiency, and data erasure requirements, this dissertation provides a comprehensive foundation for trustworthy FL systems in practice.
Date of Award7 Jan 2026
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorCong WANG (Supervisor)

Cite this

'