DOI: 10.1021/acs.jcim.5c03214 ISSN: 1549-9596

Language of Toxicity: An eXplainable Artificial Intelligence Approach

Nicola Amoroso, Ester Pantaleo, Fulvio Ciriaco, Francesca Cutropia, Nicola Gambacorta, Fabrizio Mastrolorito, Alfonso Monaco, Angelica Orfino, Roberto Bellotti, Orazio Nicolotti

Abstract

Toxicity prediction in small molecules represents a fundamental challenge in drug development and chemical safety assessment. Traditional approaches heavily rely on predefined molecular descriptors or fingerprints, potentially limiting the ability to capture complex and nonlinear structure–activity relationships. Here, we present a descriptor-free, language-inspired framework that can be applied to different toxicity prediction tasks within a unified architecture. The model proposed combines a multiscale Convolutional Neural Network (CNN) layer to capture chemical patterns at different scales and a Gated Recurrent Unit (GRU) layer to capture the sequential nature of these patterns. This architecture also exploits an attention mechanism that computes attention weights across the sequence, enabling the model to focus on the most relevant molecular substructures for toxicity prediction. Toxic and nontoxic chemicals, represented by canonical SMILES, are investigated as the words of two languages which have to be discriminated; using eight different end points, the model provided an accurate description of toxicity patterns, with an average Area under the ROC curve (AUC) of 0.83 (min: 0.70, max: 0.94) under repeated cross-validation. The models were trained on relatively small data sets (∼1000 samples) and often strongly imbalanced, two important challenges that highlight the opportunities for future improvement; moreover, the proposed attention-based framework offers a representation of the molecular regions influencing model predictions, providing a basis for future investigations into toxicity-related structural patterns and potentially supporting hypothesis generation in drug design or drug repurposing applications.