DOI: 10.3828/index.2026.30 ISSN: 0019-4131

Natural-language processing (NLP) and indexing, Part 3. Semantic search

Donald Howes

This article explores asymmetric extractive semantic search, using the large language model (LLM) SBERT. It first examines traditional exact, keyword, and fuzzy match search algorithms, discussing their features and limitations. It then explores how semantic search leverages the capacity of an LLM to capture query intent and context beyond literal word matches. An open-source application demonstrates asymmetric extractive semantic search on a public–domain PDF on Ancestral Pueblo ruins in Arizona. Sample queries on site architecture, burial practices, and areal comparisons return relevant passages, illustrating the method’s accuracy and utility to indexers. Appendices detailing the architectures of BERT and SBERT are provided.

You shall know a word by the company it keeps.

John Rupert Firth, linguist (1957)