NextTopDocker: A Large-Scale Docking-Power Benchmark Reveals Limitations of Current End-to-End Machine-Learning Docking and the Strength of Hybrid Rescoring
Cao-Minh Truong, Pedro J. Ballester, Olivier Taboureau, Viet-Khoa Tran-NguyenAbstract
Predicting three-dimensional binding orientations of drug-like molecules remains challenging in structure-based drug design. Despite methodological advances, docking performance is often assessed on small and outdated benchmarks. We present “NextTopDocker,” a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank. Developed with open-source tools, it includes crystallographic structures, Smina-generated docking poses, and ligand-similarity-aware training subsets. We benchmarked four state-of-the-art machine-learning (ML) docking frameworks (DeepDock, Interformer, SurfDock, and Uni-Mol Docking v.2) against classical (Smina) and hybrid baselines (GNINA 1.3 and logistic regression using Smina and GNINA 1.3 scores). Interformer alone matched the docking power of logistic regression on Smina poses, while the others showed dependence on downstream physics-based correction. Most raw ML-generated poses displayed steric clashes and/or implausible geometries, highlighting the need for physics-informed constraints in autonomous docking. “NextTopDocker” is available at https://github.com/caominhtr/NextTopDocker and https://zenodo.org/records/17492994.