Skip to main navigation Skip to search Skip to main content

SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models

  • Mohamed Afane*
  • , Abhishek Satyam
  • , Ke Chen
  • , Tao Li
  • , Junaid Farooq
  • , Juntao Chen*
  • *Corresponding author for this work

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Backdoor attacks create significant security threats to language models by embedding hidden triggers that manipulate model behavior during inference, presenting critical risks for AI systems deployed in healthcare and other sensitive domains. While existing defenses effectively counter obvious threats such as out-of-context trigger words and safety alignment violations, they fail against sophisticated attacks using contextually-appropriate triggers that blend seamlessly into natural language. This paper introduces three novel contextually-aware attack scenarios that exploit domain-specific knowledge and semantic plausibility: the ViralApp attack targeting social media addiction classification, the Fever attack manipulating medical diagnosis toward hypertension, and the Referral attack steering clinical recommendations. These attacks represent realistic threats where malicious actors exploit domain-specific vocabulary while maintaining semantic coherence, demonstrating how adversaries can weaponize contextual appropriateness to evade conventional detection methods. To counter both traditional and these sophisticated attacks, we present SCOUT (Saliency-based Classification Of Untrusted Tokens), a novel defense framework that identifies backdoor triggers through token-level saliency analysis rather than traditional context-based detection methods. SCOUT constructs a saliency map by measuring how the removal of individual tokens affects the model's output logits for the target label, enabling detection of both conspicuous and subtle manipulation attempts. We evaluate SCOUT on established benchmark datasets (SST-2, IMDB, AG News) against conventional attacks (BadNet, AddSent, SynBkd, StyleBkd) and our novel attacks, demonstrating that SCOUT successfully detects these sophisticated threats while preserving accuracy on clean inputs, establishing a robust defense for securing AI systems against next-generation backdoor threats. © 2025 IEEE.
Original languageEnglish
Title of host publicationProceedings - 2025 IEEE International Conference on Big Data
EditorsCheng-Zhong Xu, Leong Hou U, Xueqi Cheng, Jing Gao, Giuseppe Polese, Hong Mei, Paul Boniol, Michiaki Tatsubori, Chen Zhao, Dawei Zhou, Xiaohua Hu
PublisherIEEE
Pages6644-6652
Number of pages9
ISBN (Electronic)979-8-3315-9447-3
ISBN (Print)979-8-3315-9448-0
DOIs
Publication statusPublished - Dec 2025
Event13th IEEE International Conference on Big Data (IEEE BigData 2025) - Macau, Macao, China
Duration: 8 Dec 202511 Dec 2025
https://conferences.cis.um.edu.mo/ieeebigdata2025/

Publication series

NameProceedings of the IEEE International Conference on Big Data, BigData
ISSN (Print)2639-1589
ISSN (Electronic)2573-2978

Conference

Conference13th IEEE International Conference on Big Data (IEEE BigData 2025)
Abbreviated titleIEEE Big Data 2025
PlaceMacao, China
CityMacau
Period8/12/2511/12/25
Internet address

Funding

J. Chen acknowledges the support through Fordham AI Research Grant (FAIR) from the Fordham Office of Research.

Research Keywords

  • Backdoor Attacks
  • Clinical Language Models
  • Data Poisoning Defense
  • Healthcare AI Security

Fingerprint

Dive into the research topics of 'SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models'. Together they form a unique fingerprint.

Cite this