DOI: 10.1075/scl.129.02dov ISSN: 1388-0373

The Chinese–Spanish parallel corpus (PaCheS)

Irene Doval, Xiaomeng Wang

Abstract

Parallel corpora are a key resource for contrastive linguistics, translation studies, and natural language processing. However, for the Chinese–Spanish language pair, existing resources are often scattered and poorly integrated, domain-specific, or rely on English as a pivot language, introducing structural and stylistic biases. This paper presents the Chinese–Spanish Parallel Corpus (PaCheS),

1
a multifunctional bilingual resource designed to address these limitations. The core of PaCheS consists of approximately 15 million words from 49 contemporary literary works, ensuring high-quality, direct human translations with balanced bidirectionality. To ensure high precision in the alignment despite the structural distance between the two languages, the corpus employs a two-stage alignment workflow: automatic alignment using Bertalign followed by semantic validation via LaBSE embeddings and manual revision. Additionally, the resource is supplemented by over 2 million bisegments from institutional and web-mined sources which have undergone rigorous filtering and normalization. All materials are accessible via a web interface which supports metadata filtering and advanced queries, making PaCheS a valuable resource for research, translation, and language learning.