Large Language Models Create Hallucinations in Response to Negated Text
Jaehyung Seo, Hyeonseok Moon, Heuiseok LimLarge language models (LLMs) have achieved significant advancements in natural language processing tasks, but they remain prone to generating hallucinations—outputs that are logically inconsistent or factually incorrect. While previous research has primarily focused on hallucinations in affirmative contexts, how negated contexts and knowledge lead LLMs into hallucination has received little attention. In this paper, we demonstrate that LLMs struggle to handle negation effectively, applying affirmative knowledge inappropriately, leading to hallucinations that contradict the negated input. Through a comprehensive analysis using probing protocols, we show that negated text induces previously underexplored types of hallucinations, compromising the logical consistency and factual accuracy of LLM outputs. We employ lens observations to trace, layer by layer, how Transformer-based models encode a negated input, revealing that the layer-wise prediction probabilities for pre- and post-negation inputs remain nearly indistinguishable. Additionally, our analysis extends to Korean, a morphologically distinct language, demonstrating that negation-induced hallucinations are not issues specific to English. Our findings suggest that current LLMs, across the evaluated tasks in both English and Korean, exhibit significant areas for improvement when processing negated text, raising concerns about their reliability in real-world applications. We discuss possible ways to mitigate these hallucinations, emphasizing the need for enhanced model architectures to better understand and process negation.