Using all 25 MultiBERTs seeds, the authors measure the degree to which natural random variation yields a correlation between model quality and implicit parse accuracy (UAS) from SAS. They find no significant correlation: UAS is uncorrelated with MLM test loss (r2 = −2821) and with grammatical capabilities measured by BLIMP (r2 = −317). This null result motivates the authors' argument that correlational evidence alone is insufficient to establish SAS as essential for grammatical capabilities, and that developmental (training-time) analysis is needed.
Evidence
correlational
Key metric
r2 = −2821 (UAS vs MLM test loss); r2 = −317 (UAS vs mean BLIMP score), across 25 MultiBERTs seeds
Caveat
The r2 values as printed are far outside the standard [0,1] range for coefficient of determination, suggesting a non-standard or misreported statistic; the authors interpret these as indicating 'complete lack of significant correlation.'