IC-1399Invariant GNNs (l=0) consistently fail to distinguish k-hop identical but globally distinct geometric graphs on the k-chain task, regardless of model depth

Shih-Hsin Wang, Yung-Chang Hsu, Justin Baker, Andrea L. Bertozzi, Jack Xin, Bao Wang

SourceRethinking the Benefits of Steerable Features in 3D Equivariant Graph Neural Networks

On a synthetic k-chain classification task where two graphs differ only in the orientation of a terminal edge (making them k-hop identical but globally distinct), all invariant GNNs achieve chance-level accuracy. SchNet, DimeNet++, and SphereNet score exactly 50.0% across all chain lengths k=2,3 and depths 1-3, while CoMEt ranges from 46.5% to 59%. In contrast, equivariant GNNs with l>=1 (MACE, GVP, EquiformerV2) reach 90-100% with sufficient depth, confirming that the inability to capture geometry between local neighborhoods is the specific failure mode of invariant architectures.

Evidence
correlational
Key metric
SchNet, DimeNet++, SphereNet: 50.0 ± 0.0% test accuracy across all k and depths; CoMEt: 46.5 ± 5.0 to 59.0 ± 11.6%; MACE l=1: 100.0 ± 0.0% at k=2 layers 2-3, 100.0 ± 0.0% at k=3 layers 2-4
Caveat
The k-chain task uses only 50 graph pairs with 50/30/20 splits, so the test set is very small (10 pairs). The equivariant GNNs EGNN and CloFNet also underperform due to over-squashing, and ESCN/EquiformerV2 show abnormal performance due to equivariance errors from their spherical activation function.
Model
SchNet, DimeNet++, SphereNet, CoMEt, MACE, GVP, EGNN, CloFNet, ESCN, EquiformerV2
Concepts
Failure mode
Methods
GWL Test, IGWL test
Related findings
IC-1400
Extraction
automatic-extraction