MAJA: Multivariate Bayesian model for discovery of shared epigenetic pathways across human phenotypes
Ilse Krätschmer, Hannah M Smith, Daniel L McCartney, Elena Bernabeu, Mahdi Mahmoudi, Archie Campbell, Janie Corley, Sarah E Harris, Simon R Cox, Riccardo E Marioni, Matthew R RobinsonAbstract
Genomic measurements of DNA methylation, gene expression or protein levels, are becoming more prevalent and are increasingly used to study health outcomes. However, most proposed association testing methods consider only marginal effects of each feature on a single outcome variable and are not set up to handle highly correlated, continuous data. Here, we introduce MAJA, a method to learn shared and outcome-specific effects for multiple traits in multi-omics data. MAJA determines the unique contribution of individual loci, genes, or molecular pathways, to variation in one or more traits, conditional on all other measured "omics" data genome-wide. Simulations show MAJA accurately finds shared and distinct associations between omics-data and multiple traits and estimates omics-specific (co)variances, allowing for sparsity and correlations within the data. Applying MAJA to 12 outcome traits in Generation Scotland methylation data (n = 18,264), we find novel shared epigenetic probes among cholesterol metabolism, osteoarthritis, blood pressure and asthma. In contrast to marginal testing, we find only 10 CpG probes with significant effects above the genome-wide background. This highlights the need for joint association testing in highly correlated methylation data from whole blood and for studies of increased sample size in order to refine epigenomic associations in observational data.