Abstract
<jats:p>Molecular docking and co-folding engines are widely used to prioritize compounds for wet-lab validation, yet their accuracy is known to vary substantially across protein targets for reasons that remain only qualitatively understood. Here we benchmark six docking and co-folding engines (RevDock, DiffDock, Boltz2, AutoDock-GPU, rDock, and PandaDock) across 14 protein families, evaluating scoring power, ranking power, docking power, and physical validity. Rather than treating engine performance as protein-family-specific, we classify all 14 families into six mechanistic groups according to which of four scoring-function simplifications, rigid receptor, pairwise additivity, fixed point charges, and implicit solvent, is most severely stressed by that family's binding site. This framework helps explain, rather than simply describe, where each engine succeeds or fails: RevDock's CNN rescoring layer mitigates the pairwise additivity and fixed-charge limitations relative to physics-only scoring, achieving the highest overall pose accuracy (73.3% of poses ≤ 2.0 A RMSD), while Boltz2's sequence-based co-folding bypasses the rigid-receptor assumption and achieves comparable affinity correlation (mean Pearson r ≈ 0.60 for both engines). PandaDock, run with expanded conformational sampling, matches RevDock on pose accuracy (72.1% of poses ≤ 2.0 A, lowest median RMSD at 0.96 A) and exceeds AutoDock-GPU on affinity correlation (mean r = 0.460), indicating that the performance of a physics-based scoring function is limited as much by search adequacy as by the scoring function itself. These results suggest that engine selection for a docking or co-folding campaign should be guided less by an engine's aggregate benchmark ranking and more by which of these four structural and physical characteristics dominate the target of interest.</jats:p>