Discourse Structure as an Interpretable Signal for Detecting Hallucinated Chain-of-Thought Reasoning in Large Language Models
Boris GalitskyLarge language models can generate fluent chain-of-thought (CoT) reasoning that appears coherent while exhibiting systematic distortions in evidence weighting and hypothesis comparison. This paper studies hallucinated CoT as a discourse-structural phenomenon, not only a factual one. We introduce a diagnostic reasoning benchmark with paired grounded and hallucinated explanations, where traces differ in how they organize evidence, alternatives, and defeaters. We extract discourse tree features that summarize evidence allocation, contrast preservation, commitment timing, and evidence integration, and combine them with the Joint Knowledge–Reasoning Hallucination Measure (JKRHM). Experiments on the synthetic diagnostic dataset and preliminary external validation on HaluBench suggest that discourse structure provides an interpretable signal for detecting reasoning hallucinations and complements existing factuality and uncertainty-based hallucination detectors. Because the HaluBench reasoning rationales are generated as an intermediate representation, these results should not be interpreted as definitive external proof of generalization. The results support a cautious conclusion: discourse analysis does not replace factual verification, but it helps expose reasoning paths that are structurally unsupported even when they are fluent and persuasive.