The architecture employs E2Former-V2 equivariant Transformer, satisfying rotational equivariance, non-covalent distance receptive fields, and single-GPU hardware efficiency. Through SO(3)→SO(2) basis transformation, node-centric Wigner-6j decomposition, and Triton streaming attention, near-linear computational complexity is achieved. The backbone stacks four layers with mixed cutoff radii: three layers of 5Å short-range attention covering all atoms, and a fourth layer at 8Å excluding hydrogen, yielding an effective receptive field of about 23Å.
Training data comprises approximately 160 million quantum chemistry labels, from OMol25 (about 140 million) and self-built UBio-Mol26 (about 19 million). UBio-Mol26 enriches methylene and amide groups, with atom-pair distances extending beyond 5-6Å, covering biomolecular features. About 64,000 periodic condensed-phase configurations are used to calibrate thermodynamic properties such as water density.
Three-stage curriculum learning totals about 1000 GPU-days: the first stage uses independent force heads, the second switches to automatic differentiation (F=-∇E), and the third introduces biomolecular data, supervising only force vectors to eliminate DFT systematic energy offsets.