DOI: 10.1075/scl.129.02dov ISSN: 1388-0373
The Chinese–Spanish parallel corpus (PaCheS)
Irene Doval, Xiaomeng WangAbstract
Parallel corpora are a key resource for contrastive linguistics, translation studies, and natural language processing. However, for the Chinese–Spanish language pair, existing resources are often scattered and poorly integrated, domain-specific, or rely on English as a pivot language, introducing structural and stylistic biases. This paper presents the Chinese–Spanish Parallel Corpus (PaCheS),
1
a multifunctional bilingual resource
designed to address these limitations. The core of PaCheS consists of approximately 15 million words from 49 contemporary
literary works, ensuring high-quality, direct human translations with balanced bidirectionality. To ensure high precision in
the alignment despite the structural distance between the two languages, the corpus employs a two-stage alignment workflow:
automatic alignment using Bertalign followed by semantic validation via LaBSE embeddings and manual revision. Additionally,
the resource is supplemented by over 2 million bisegments from institutional and web-mined sources which have undergone
rigorous filtering and normalization. All materials are accessible via a web interface which supports metadata filtering and
advanced queries, making PaCheS a valuable resource for research, translation, and language learning.