Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Endless Jailbreaks with Bijection Learning
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-572
Bijection learning achieves state-of-the-art jailbreak ASR on frontier models, with peak ASR increasing with model capability
IC-573
Model capabilities on MMLU degrade monotonically as bijection encoding complexity increases
IC-574
Guard models fail to effectively mitigate bijection attacks even at capability parity with the target model