IC-573Model capabilities on MMLU degrade monotonically as bijection encoding complexity increases

Brian R.Y. Huang, Maximilian Li, Leonard Tang

SourceEndless Jailbreaks with Bijection Learning

To test the computational overload hypothesis, the authors evaluate four models on 10-shot MMLU while the questions and answer labels are encoded in bijection languages of varying dispersion and encoding length. For Claude 3 Haiku, Claude 3.5 Sonnet, GPT-4o-mini, and GPT-4o, MMLU accuracy decreases monotonically as dispersion and encoding length increase. This degradation is the proposed mechanism by which bijection attacks bypass safety training: the translation task consumes computational resources that would otherwise be available for safety classification.

Evidence
correlational
Caveat
MMLU results are reported only in figures (Figures 6 and 11); no specific accuracy numbers are printed in the text. The evaluation uses 10-shot in-context examples of MMLU questions in bijection language, which may not perfectly isolate the translation overhead.
Model
Claude 3 Haiku, Claude 3.5 Sonnet, GPT-4o mini
Concepts
Failure mode
Datasets
MMLU / MMLU-Math [eval]
Related findings
IC-572, IC-574
Extraction
automatic-extraction