IC-926Across 25 MultiBERTs seeds, UAS does not correlate with MLM test loss or BLIMP accuracy

Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L Leavitt, Naomi Saphra

SourceSudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs

Using all 25 MultiBERTs seeds, the authors measure the degree to which natural random variation yields a correlation between model quality and implicit parse accuracy (UAS) from SAS. They find no significant correlation: UAS is uncorrelated with MLM test loss (r2 = −2821) and with grammatical capabilities measured by BLIMP (r2 = −317). This null result motivates the authors' argument that correlational evidence alone is insufficient to establish SAS as essential for grammatical capabilities, and that developmental (training-time) analysis is needed.

Evidence
correlational
Key metric
r2 = −2821 (UAS vs MLM test loss); r2 = −317 (UAS vs mean BLIMP score), across 25 MultiBERTs seeds
Caveat
The r2 values as printed are far outside the standard [0,1] range for coefficient of determination, suggesting a non-standard or misreported statistic; the authors interpret these as indicating 'complete lack of significant correlation.'
Model
MultiBERTs
Datasets
BLIMP [eval], Penn Treebank / Penn Treebank 3 [eval]
Methods
SAS Probe [primary], MLM Scoring [eval]
Related findings
IC-925
Extraction
automatic-extraction