Genome-Wide Association Study and Genomic Selection for Average Daily Gain in Ashidan Yak
Zhicheng Wang, Xiaoming Ma, Guangwei Hu, Jianwu Jing, Yongfu La, Wenwen Ren, Baicheng Zhou, Hongkang Li, Min Chu, Xiaoyun Wu, Ping Yan, Xian Guo, Chunnian LiangAverage daily gain (ADG) is a core quantitative trait determining the economic benefits of Ashidan yak, a polled new breed adapted to cold barn feeding on the Qinghai–Tibet Plateau. Unraveling its complex genetic architecture is crucial for early molecular breeding selection. In this study, high-depth whole-genome resequencing (WGS) data from 474 Ashidan yaks were used to conduct combined evaluation of genome-wide association study (GWAS) and genomic selection (GS). During GWAS analysis, sex and measurement batch were included as fixed effects, while birth weight and principal components (PC1–PC3) were incorporated as covariates. Multi-model association analysis using GLM, MLM, and FarmCPU was performed on 3.36 million LD-pruned SNPs. The genomic inflation factors (λ ≈ 1.0) for MLM and FarmCPU confirmed effective elimination of population stratification. A total of 11 genome-wide significant SNP loci and 7 key candidate genes including PDE10A, RAD51B, BCAS3 and KCNH8 were identified via the FarmCPU model. Functional enrichment analysis indicated that these gene clusters are significantly involved in cAMP signaling pathway, regulation of ion channel activity, as well as extracellular matrix remodeling of blood vessels and skeletal muscle cells. For genomic selection, a single-trait GBLUP model was constructed using 22.87 million high-density raw SNPs to fully capture polygenic minor effects. Moderately high narrow-sense genomic heritability of ADG was estimated at h2 = 0.3233 (p < 0.05). The average independent prediction accuracy across the 10-fold cross-validation reached an average of R = 0.14 ± 0.07. To evaluate marker prioritized genomic evaluation without data leakage, a strict 10-fold cross-validation scheme was implemented, where the top 1% high-priority variant set (~228,000 SNPs) was screened independently within each training fold. The resulting unbiased prediction accuracy reached R = 0.1328 ± 0.1566 (with an average RMSE of 0.3793 ± 0.0066 and a regression slope of 0.4183 ± 0.5055). Comparing this with the unselected whole-genome baseline (R = 0.14 ± 0.07) indicates that naive marker selection based solely on GBLUP effect size in small reference cohorts is influenced by sampling variance, highlighting the need to integrate multi-omics functional annotations for future custom breeding array development. This study provides quantitative insights into the polygenic architecture of ADG in yaks, offering baseline data for genomic selection and custom array development for indigenous livestock on the Qinghai–Tibet Plateau.