Local Geometry Recovers but Cooperative Structure Does Not: Residue-Resolved Limits of All-Atom Reconstruction from Single-Bead Coarse-Grained Disordered Protein Ensembles
Jianxiang Huang, Xin Qiao, Ning Liu, Zongtao Chai, Shaoyong LuAbstract
Intrinsically disordered proteins (IDPs) drive diverse cellular processes through broad conformational ensembles, but experimental characterization of these ensembles is sparse and high-quality training data is scarce, holding back artificial intelligence (AI) approaches to ensemble prediction. Single-bead coarse-grained (CG) force fields such as CALVADOS, parametrized directly against experimental observables, currently provide a more reliable route to disordered ensembles than direct AI prediction. CG sampling lacks atomistic resolution and must be paired with backmapping; the atomistic information recoverable from this two-step process is shaped jointly by the CG representation and the backmapping algorithm, and the interplay between these contributions is not well characterized. We benchmarked CODLAD, a recently published latent-diffusion backmapping pipeline, on 23 Protein Ensemble Database (PED) systems using a dual-input design: the same architecture receives either PED-reference Cα coordinates or independently sampled CALVADOS Cα trajectories. This design separates paired reconstruction error, measurable for PED+CODLAD, from unpaired CG-input-associated ensemble deviations, measurable for CG+CODLAD at the distribution level. Reconstruction from PED conformers achieved 0.56 ± 0.08 Å backbone root-mean-square deviation and preserved local geometry. Reconstruction from CALVADOS-sampled Cα trajectories preserved the Cα framework and bond geometry, while ensemble-level deviations relative to PED were localized to proline backbone geometry and cooperative secondary structure, both consistent with information not carried by an unconstrained single-bead representation. The benchmark quantifies which atomistic properties can be recovered after projecting a CG ensemble into all-atom space and identifies improved CG geometric encoding and sequence-conditioned AI priors as the directions for further progress.