DOI: 10.1515/flin-2026-0052 ISSN: 0165-4004

Are complex sentences reliable proxies for text-level linguistic complexity? A dependency-based corpus analysis of English

Jinlu Liu, Jinyi Zhang, Gaiying Chai

Abstract

This study investigates whether English complex sentences can serve as reliable proxies for the overall linguistic complexity of the texts in which they appear. Focusing on both lexical and syntactic complexity, we analyze a dependency-annotated version of the Brown Corpus (1 million words, 15 genres), comparing full-text measures against those derived exclusively from complex sentences. Lexical complexity is operationalized via Standardized Type-Token Ratio (STTR) and Lexical Density (LD); syntactic complexity is measured using Mean Dependency Distance (MDD) for linear complexity and Mean Hierarchical Distance (MHD) for hierarchical complexity. Results show significant positive correlations between complex sentences and their host texts across all four indices and across all genres. Regression analyses further indicate that complex-sentence-based measures strongly predict full-text LD and MHD (R 2  > 0.8, large effect sizes), while their predictiveness is moderate for STTR and MDD (R 2  ≈ 0.42, medium-to-large effect sizes). These findings demonstrate that complex sentences capture key dimensions of textual complexity, with LD and MHD as robust proxies and STTR and MDD as more moderate ones. This offers a linguistically motivated and computationally efficient entry point for large-scale text analysis, provided that index-specific representativeness is taken into account. The study contributes to corpus linguistics by empirically validating a reductionist yet theoretically grounded approach to complexity measurement, and opens new avenues for research on the representativeness of syntactic structures in text-level analysis.