SourceEndless Jailbreaks with Bijection Learning
To test the computational overload hypothesis, the authors evaluate four models on 10-shot MMLU while the questions and answer labels are encoded in bijection languages of varying dispersion and encoding length. For Claude 3 Haiku, Claude 3.5 Sonnet, GPT-4o-mini, and GPT-4o, MMLU accuracy decreases monotonically as dispersion and encoding length increase. This degradation is the proposed mechanism by which bijection attacks bypass safety training: the translation task consumes computational resources that would otherwise be available for safety classification.