IC-268ESM-2 and ProGen-2 zero-shot fitness prediction follows an inverted U-shape as a function of wild type sequence likelihood, with both under- and over-preferred sequences degrading performance

Cade W Gordon, Amy X. Lu, Pieter Abbeel

SourceProtein Language Model Fitness is a Matter of Preference

Across 217 deep mutational scan studies from ProteinGym, the pseudo log likelihood (PLL) of the wild type sequence predicts the Spearman correlation of zero-shot mutation effect predictions for both ESM-2 and ProGen-2 at multiple scales. Performance improves as likelihood increases from very low values, but degrades again at very high likelihoods, producing a concave-down parabola. This effect is most pronounced in larger models (ESM-2 15B, ProGen-2 XL), where sequences with probabilities near 1 show the greatest degradation. Additionally, the explanatory power of likelihood for performance decreases as model size increases, an inverse scaling law.

Evidence
correlational
Key metric
Second-order polynomial regression on PLL vs DMS Spearman yields concave-down parabola for all model sizes; ESM-2 performance degrades past 650M parameters on ProteinGym task averages; effect magnified for sequences with probabilities near 1 (log scale 0) in ESM-2 15B and ProGen-2 XL
Caveat
The linear additivity assumption of the log-odds-ratio fitness metric cannot model epistatic interactions, and mutants must be identical in length to the wild type.
Model
ESM-2, ProGen-2
Concepts
Failure mode
Datasets
ProteinGym [eval]
Methods
Spearman rank correlation [eval]
Related findings
IC-269, IC-270
Extraction
automatic-extraction