DOI: 10.1158/1538-7445.pediatric26-a008 ISSN: 0008-5472

Abstract A008: OV-BENCH: a 100-task benchmark for evaluating AI models across the oncolytic virus design pipeline for pediatric CNS malignancies

Archit Kalra, Avinash Valuveri

Abstract

Pediatric CNS malignancies including high-grade glioma, medulloblastoma, and diffuse midline glioma remain among the deadliest childhood cancers. Oncolytic virotherapy has emerged as a promising immunotherapeutic modality, with several agents in early-phase pediatric trials. However, the design space for oncolytic viruses is vast, encompassing genome engineering, transgene payload selection, capsid retargeting, and safety assessment. Despite the emergence of powerful genomic language models, no standardized benchmark exists to evaluate AI capabilities across the full oncolytic virus design pipeline. We developed OV-BENCH, a 100-task benchmark spanning eight categories: viral genome generation, transgene design, capsid and structural protein design, immunogenicity prediction, viral genome understanding, tumor targeting, safety assessment, and integrative end-to-end design. Each task specifies public datasets, evaluation metrics, and biological rationale. The benchmark references 16 models and 13 public databases, with quality control incorporating CheckV and GeNomad following methodologies from recent bacteriophage genome design work. Preliminary experiments evaluated Evo 2 (40B) zero-shot on oncolytic virus chassis despite eukaryotic viruses being excluded from its training data. In perplexity scoring across seven genomes, Evo 2 assigned oncolytic virus chassis (HSV-1: 3.40, adenovirus 5: 3.57, vaccinia: 3.22) significantly lower perplexity than random DNA (3.81), demonstrating transferable DNA grammar learning, while in-training organisms scored lower still (lambda phage: 1.12, E. coli: 1.27). In a gene essentiality prediction task on HSV-1, Evo 2 zero-shot log-likelihood features predicted gene essentiality with 73.8% leave-one-out cross-validation accuracy (permutation test p=0.002), correctly identifying four of five clinically validated oncolytic virus deletion targets (UL23/thymidine kinase, US12/ICP47, US3, US11) as non-essential. These results establish that genomic language models encode biologically meaningful signals for oncolytic virus engineering, motivating systematic benchmarking to accelerate computational OV design for pediatric cancers.

Citation Format:

Archit Kalra, Avinash Valuveri. V-BENCH: a 100-task benchmark for evaluating AI models across the oncolytic virus design pipeline for pediatric CNS malignancies [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Bridging Discovery and Clinical Impact in Pediatric Cancer; 2026 Sep 22-25; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2026;86(18_Suppl_1):Abstract nr A008.