Y Chromosome Haplogroups and Susceptibility to Esophageal Squamous Cell Carcinoma: An Exploratory Association and Diagnostic Modeling Study
Xueji Shi, Nian Wang, Wenjing Zhou, Huiyu Zhong, Tangyuheng Liu, Mengyuan Song, Yi Zhou, Xingbo SongABSTRACT
Background
Esophageal squamous cell carcinoma (ESCC) showed remarkable male predominance, but the potential relevance of Y chromosome background remains unclear. This exploratory study investigated Y chromosome haplogroup distribution and its possible contribution to ESCC diagnostic modeling.
Methods
We recruited 517 male ESCC patients and 447 age‐matched healthy men from Sichuan, China. Y chromosomal short tandem repeats (Y‐STRs) were used to assess paternal lineage structure and exclude cryptic relations, and Y chromosomal single nucleotide polymorphisms (Y‐SNPs) were used to classify Y chromosome haplogroups. We analyzed the association between Y chromosome haplogroups and genetic susceptibility to ESCC by false discovery rate (FDR) correction and adjusted logistic regression. Repeated bootstrap‐Boruta and Boruta analyses were performed to evaluate feature selection stability of haplogroups. Nomogram model with and without O2a1b haplogroup was constructed using clinical variables, and compared by receiver operating characteristic (ROC), calibration, decision curve analyses, and DeLong test. Exploratory RNA‐seq was performed in tissue and blood samples from O2a1b‐positive and non‐O2a1b ESCC patients.
Results
Principal component analysis (PCA) and median‐joining tree showed substantial overlap between ESCC cases and controls. O2a1b had nominal positive association with ESCC, whereas O1b1a showed a nominal negative association with ESCC. However, no haplogroup remained significant after FDR correction. In adjusted logistic regression, O2a1b remained significant with ESCC, but repeated bootstrap‐Boruta analysis showed low feature‐selection stability for O2a1b and O1b1a. The nomogram without haplogroup, composed of BMI, CEA, ALB, ALP, TBA, and NLR, achieved AUCs of 0.883 and 0.888 in the training and validation cohort, respectively. Adding O2a1b did not significantly improve the performance of the nomogram. RNA‐seq suggested potential involving pathways, such as “Hippo signaling pathway” and “focal adhesion signaling pathway.”
Conclusion
Y chromosome haplogroups were not recognized as robust markers for ESCC diagnosis in our cohort. Although these findings should be regarded as exploratory and hypothesis‐generating, this study undermined the potential value of Y chromosome haplogroups and paternal lineage in male ESCC susceptibility. Further validation in larger independent cohorts is required.