DOI: 10.3390/e28101083 ISSN: 1099-4300

The Distribution of Lexical Information Within Dependency Spans: Cross-Linguistic Evidence from 22 Languages

Xiangyu Kang, Shuiyuan Yu

Dependency distance measures the linear separation between syntactically related words but does not describe how information is distributed across their interval. We analyzed summed XGLM-2.9B word surprisal in 1,346,122 dependency spans from 22 Universal Dependencies languages at distances 4–10. Clustering 154 position-aligned profiles yielded a broadly descending type and a type with a terminal rise; 20 languages retained their assignment across distances. Leave-one-language-out comparison did not support a common three-segment structure: the one-standard-error rule selected two segments for the descending type and three for the rising type. The descending type’s initial decline disappeared after separate adjustments for sentence position and preceding context. Standardizing word-length and model subword-count distributions reduced the rising type’s mean final-word increase from +1.996 to +0.408 bits (95% interval [−0.024, +0.898]). In held-out sentences, real dependencies in eight languages independently classified by a training-set terminal decrease by 0.928 bits more than nearby nonedge intervals sharing the final part-of-speech tag (95% interval [−1.498, −0.217]). A shared part-of-speech component was negatively associated with terminal surprisal in the original descending type. Dependency-span alignment reveals recurring information profiles and locates a conditional terminal contrast for further tests of syntactic prediction.