DOI: 10.1162/tacl.a.805 ISSN: 2307-387X

Friend or Foe: How LLMs’ Safety Mind Gets Fooled by Intent Shift Attack

Peng Ding, Jun Kuang, Wen Sun, Zongyu Wang, Xuezhi Cao, Xunliang Cai, Jiajun Chen, Shujian Huang

Abstract

Large language models (LLMs) remain vulnerable to jailbreaking attacks despite their impressive capabilities. Investigating these vulnerabilities is crucial for developing robust safety mechanisms. Existing attacks primarily bypass LLM safeguards by introducing additional context or adversarial tokens, while leaving the core harmful intent intact. In this paper, we introduce ISA (Intent Shift Attack), which conceals malicious intent from LLMs through subtle edits. More specifically, we establish a taxonomy of intent transformations and leverage them to generate attacks that may be misperceived by LLMs as benign requests. Unlike prior methods relying on complex tokens or lengthy context, our approach requires only limited modifications to the original request, yielding natural, fluent, and seemingly harmless prompts. Extensive experiments on both open-source and commercial LLMs demonstrate that ISA effectively induces safety failures, achieving high attack success rates through simple linguistic transformations. More critically, fine-tuning models solely on benign data reformulated with ISA templates elevates the success rates to nearly 100%. For defense, we evaluate existing methods and demonstrate their limited effectiveness against ISA, while exploring both training-free and training-based mitigation strategies. Our findings reveal fundamental challenges in intent inference for LLM safety and underscore the need for more effective defenses. Our code and datasets are available at https://github.com/NJUNLP/ISA.