DOI: 10.3390/jintelligence14100234 ISSN: 2079-3200

Opening the Window of the Mind: Do Large Language Models Possess Human-Like Visual Attention?

Houji Jin, Ganyu Gui, Ziyi Wang, Zhicheng Cai, Zhijie Tang, Jiayi Zhang, Jian Wang, Huanghao Feng

When people read a spatial description, their eyes move through empty space as if the described scene were present, a signature of imagery-driven attention documented by Spivey and Geng. Language models have no eyes, but they are routinely credited with spatial understanding on the basis of verbal answers, and such answers are ambiguous: a model can name the right direction by echoing a directional word the passage happens to contain. We ask whether responses to spatial descriptions are governed by the configuration described or by the vocabulary used to describe it. Adapting the Spivey and Geng materials, we hold each scene’s layout fixed and vary its linguistic support across four versions: one naming the direction, one conveying only the configuration, one reordering the same sentences, and one inserting a directional word that contradicts the layout. Five models from the Qwen and GLM families were probed with a forced-choice direction report, an unconstrained narration, a two-dimensional reconstruction, and a plain relational question, together with control conditions that vary the position and register of the inserted word, shuffle the object list in the prompt, remove the imagery framing, and offer an explicit indeterminate option. Analysis treats model-by-scene cells as units, with cluster-bootstrap intervals and mixed-effects models, across 14,900 responses. Passages that convey a layout without naming a direction are answered near ceiling, reconstructed and narrated in the correct object order, and unaffected by reordering, by shuffling the prompt’s object list, or by removing the imagery framing: directional vocabulary is not what these responses track. Three apparent limits on that competence proved to be artifacts of measurement. A misleading word is resisted on three quarters of trials when it precedes the layout and on none when it follows it, so the manipulation indexes recency rather than representational strength; a downward bias on axis-free scenes disappears once declining is permitted; and narrations judged unfaithful by direction coding are faithful when scored by object order. Layout recovery is robust, and spatial evaluations that vary a single response format risk measuring their instrument.