DOI: 10.1002/ppp3.70266 ISSN: 2572-2611

A Scalable Large Language Model‐Aided Approach for Mining Plant‐Compound Pairs From Scientific Literature

Adam Richard‐Bollans, Francesco Civita

Societal Impact Statement

Understanding plant chemicals is vital for medicine and agriculture, but global data remain fragmented and biased. We developed an AI‐powered tool to automatically extract plant‐compound information from scientific literature at scale and low cost. Our analysis revealed significant knowledge gaps, particularly in the Southern Hemisphere. By efficiently uncovering thousands of potential new plant‐chemical links, this work provides a crucial resource for bioprospecting, taxonomy and conservation efforts, helping to ensure the benefits of nature's chemical diversity are fully realised.

Summary

An enormous amount of phytochemical knowledge is contained within published scientific literature, much of which is not captured in existing phytochemical datasets. This study focuses on providing a robust and scalable approach to unlock these data by employing large language models combined with automated and manual verification.

We analysed data from two major phytochemical datasets, WikiData and KNApSAcK, highlighting potential taxonomic and geographic biases. A pipeline was developed using the DeepSeek‐V3 large language model to extract phytochemical occurrences from literature before undergoing automated filtering through standardization of plant and compound names. The approach is evaluated against a subset of records from WikiData, through manual verification and against a subset of literature known to not contain phytochemical occurrences.

The study shows that the DeepSeek model achieves high precision in extracting plant‐compound pairs from text. Application to 257 phytochemistry papers uncovered 7,447 potential pairs, 4,832 of which are not found in WikiData or KNApSAcK. Targeted application of our approach to phytochemistry publications related to species native to Colombia, which is under‐represented in WikiData and KNApSAcK, reveals phytochemical data for four species not previously recorded in these phytochemical datasets.

This work demonstrates that large language models like DeepSeek offer a scalable, cost‐effective solution for mining phytochemical data. By automating extraction from literature, this approach can efficiently identify thousands of novel plant‐compound occurrences, helping to address geographic and taxonomic biases in existing databases and significantly expand our knowledge of global plant chemical diversity.