DOI: 10.1371/journal.pone.0351281 ISSN: 1932-6203

Arabic text preprocessing for dynamic searchable symmetric encryption: An empirical evaluation and preprocessor selection framework

Kholoud Saad Al-Saleh

Dynamic Searchable Symmetric Encryption (DSSE) enables keyword search over encrypted data without revealing plaintext to the server. Arabic morphological richness — where a single root generates dozens of surface forms — creates substantial challenges for encrypted search: unprocessed vocabularies inflate encrypted-index size and transmission cost, and limit retrieval recall by failing to match morphological variants. This paper presents the first empirically grounded preprocessor selection framework for Arabic DSSE, derived from a systematic evaluation of four Arabic preprocessing strategies — the Khoja stemmer, ISRI stemmer, Lucene Arabic Analyzer, and Farasa segmenter — against a normalization-only baseline within the Incidence Matrix DSSE (IM-DSSE) scheme on the Khaleej Arabic news corpus. We evaluate vocabulary size, search latency, search quality, and retrieval breadth on an annotated 1,400-document corpus using 56 benchmark queries, and assess scalability on cloud infrastructure across corpus sizes up to 45,500 documents. Stemming reduces vocabulary by up to 83% relative to the normalization-only baseline, proportionally reducing encrypted-index size and transmission cost. We find that per-query search latency is driven primarily by result-set size rather than vocabulary size: the normalization-only baseline is fastest per query because it matches the fewest documents, while preprocessors that broaden retrieval incur higher latency. At scale, all five configurations — including the normalization-only baseline — operate successfully to 45,500 documents, and the result-set-size effect on latency persists. In terms of search quality, light stemming and morphological segmentation achieve the best trade-off between retrieval precision and recall, while root-based stemmers sacrifice precision without commensurate recall gains. Based on these findings, the framework provides actionable, evidence-based guidelines mapping deployment requirements — memory scalability, search latency, search quality, and retrieval breadth — to concrete preprocessor choices for Arabic DSSE system designers.