DOI: 10.3390/electronics15163677 ISSN: 2079-9292

Recent Advances and Open Challenges in Mitigating Inference-Time Attacks on Large Language Models

Berkay Özçam, Mustafa Kara, Muhammed Ali Aydın, Hasan Hüseyin Balık

The rapid integration of Large Language Models into high-stakes domains has elevated inference-time attacks into a primary security concern for production deployments. These attacks are adversarial techniques that exploit models exclusively through their input–output interface. The existing survey literature lacks a dedicated and structured treatment that jointly maps the attack surface and systematically evaluates the mitigation strategies developed against it. This paper addresses this gap through two original taxonomic contributions. First, LLM vulnerabilities are organized into a three-layer attack surface taxonomy stratified by lifecycle stage, establishing the theoretical primacy of the inference time category. Second, to directly address how these attacks can be mitigated, a defense taxonomy spanning three axes, namely prompt-level, inference-time, and training-time interventions, is proposed, within which 30 mitigation mechanisms published from 2024 onwards are systematically analyzed. Building on this taxonomy, an intersectional comparative analysis is conducted across three dimensions: defense-attack coverage, security-utility-latency tradeoffs, and white-box versus black-box applicability, in order to evaluate how effectively current mitigation strategies neutralize each attack category. These dimensions are further synthesized into a practitioner decision framework that maps deployment constraints to concrete defense configurations and identifies two structural coverage gaps that persist regardless of access level or latency budget. The resulting Defense-Attack Coverage Matrix demonstrates that no single defense mechanism provides comprehensive protection, and that robust deployment mandates layered, complementary strategies. The analysis further reveals that the fundamental unresolved tension limiting effective mitigation is the trade-off between adversarial robustness and model utility, with over-refusal and capability degradation constituting the primary practical barriers to deploying these defenses. Finally, open challenges related to multimodal attack surfaces, agentic LLM security, and the absence of standardized evaluation frameworks are identified, together with concrete future research directions. The taxonomies and analyses presented are intended to serve as an actionable reference for both researchers and practitioners tasked with mitigating inference-time attacks in secure LLM deployments.

More from our Archive