Skip to main navigation Skip to search Skip to main content

Uncovering Gradient Inversion Risks in Practical Language Model Training

  • Xinguo Feng
  • , Zhongkui Ma
  • , Zihan Wang
  • , Eu Joe Chegne
  • , Mengyao Ma
  • , Alsharif Abuadbba
  • , Guangdong Bai*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

The gradient inversion attack has been demonstrated as a significant privacy threat to federated learning (FL), particularly in continuous domains such as vision models. In contrast, it is often considered less effective or highly dependent on impractical training settings when applied to language models, due to the challenges posed by the discrete nature of tokens in text data. As a result, its potential privacy threats remain largely underestimated, despite FL being an emerging training method for language models. In this work, we propose a domain-specific gradient inversion attack named Grab (gradient inversion with hybrid optimization). Grab features two alternating optimization processes to address the challenges caused by practical training settings, including a simultaneous optimization on dropout masks between layers for improved token recovery and a discrete optimization for effective token sequencing. Grab can recover a significant portion (up to 92.9% recovery rate) of the private training data, outperforming the attack strategy of utilizing discrete optimization with an auxiliary model by notable improvements of up to 28.9% recovery rate in benchmark settings and 48.5% recovery rate in practical settings. Grab provides a valuable step forward in understanding this privacy threat in the emerging FL training mode of language models. © 2024 Copyright held by the owner/author(s).
Original languageEnglish
Title of host publicationCCS '24
Subtitle of host publicationProceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security
PublisherAssociation for Computing Machinery
Pages3525-3539
Number of pages16
ISBN (Print)9798400706363
DOIs
Publication statusPublished - Dec 2024
Externally publishedYes
Event31st ACM SIGSAC Conference on Computer and Communications Security (CCS 2024) - Salt Lake City, United States
Duration: 14 Oct 202418 Oct 2024

Publication series

NameCCS - Proceedings of the ACM SIGSAC Conference on Computer and Communications Security

Conference

Conference31st ACM SIGSAC Conference on Computer and Communications Security (CCS 2024)
PlaceUnited States
CitySalt Lake City
Period14/10/2418/10/24

Funding

We thank our anonymous shepherd and reviewers for their constructive comments. This work is partially supported by Australian Research Council Discovery Projects (DP230101196, DP240103068).

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 4 - Quality Education
    SDG 4 Quality Education

Research Keywords

  • Federated Learning
  • Gradient Inversion
  • Language Models

Fingerprint

Dive into the research topics of 'Uncovering Gradient Inversion Risks in Practical Language Model Training'. Together they form a unique fingerprint.

Cite this