DOI: 10.3390/e28080869 ISSN: 1099-4300

Everything Is Prediction: Modern Machine Learning as Bayesian Inference

Nicholas G. Polson, Vadim Sokolov, Refik Soyer

We argue that the core methods of modern machine learning—conformal prediction, large language models and in-context learning, and generative/diffusion models—are not rivals to Bayesian inference but implementations of it, almost always of its predictive (de Finetti) form rather than its parameter-centric (prior-to-posterior) form. (1) Background: A quarter-century after Breiman contrasted the “data-modeling” and “algorithmic” cultures of statistics, we revisit that dichotomy and argue that the predictive view dissolves it—both cultures target the one-step-ahead density p(yn+1∣y1:n), differing only in how they compute it. (2) Methods: We organize the modern toolkit around this predictive object and around amortization—replacing per-dataset inference with a single map learned by simulation—using Generative Bayesian Computation (GBC) as the connective spine. (3) Results: Conformal prediction, autoregressive language models, prior-data fitted networks, and score-based diffusion are each shown to construct, calibrate, or sample from the predictive object, summarized in a single “Rosetta” table; because all are fit by proper scoring rules—equivalently, by Bregman divergences—the information-theoretic frame is the natural unifier. (4) Conclusions: The equivalence is exact in idealized limits, asymptotic under exchangeability and its martingale relaxations, and measurably approximate for trained models—a three-grade taxonomy that we make explicit, row by row, and that bounds the thesis: prediction is not attribution, and the predictive view is deliberately silent about causal structures.

More from our Archive