DOI: 10.1093/bioadv/vbag238 ISSN: 2635-0041

Inductive bias influences the spatial scale of biological features learned from images

Jacob I Evarts, Jason Y Cain, Po-Hao Chiu, Ian Jan, Nancy L Allbritton, Neda Bagheri

Abstract

The inductive bias of a deep learning model influences the features it extracts from biological images, making model selection a critical scientific decision. We systematically compare representations learned from scratch without pre-training by convolutional neural networks (CNNs), Vision Transformers (ViTs), and Fourier neural operators (FNOs)—an emerging architecture largely unexplored in biological imaging—to determine how their distinct learning mechanisms shape feature learning from spatiotemporal data. Using a self-supervised representation learning framework on microscopy images of two-dimensional (2D) gastruloids, as well as a distinct synthetic tumor image dataset, we show that while CNNs and FNOs achieve comparable accuracy on biological tasks, they learn features at different scales. Visual interpretation methods reveal that CNNs prioritize local, fine-grained details, whereas FNOs capture global, low-frequency structures, in line with overall population morphology. These findings demonstrate that deep learning model architecture is a choice that shapes the biological scale of extracted features, highlighting the need to align a model’s inductive bias with the scientific question. All source code for the representation learning model, the tumor simulation dataset, and the gastruloid microscopy imaging dataset are available at the corresponding DOIs: 10.5281/zenodo.19672838, 10.5281/zenodo.19361370, 10.5281/zenodo.19373112, respectively.

More from our Archive