IC-011Jailbreaking LLMs can reduce refusal rates and improve alignment with human preferences
Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, Bernhard Schölkopf
The paper applies an uncensoring technique to four open-source LLMs (Llama 3.1 8B, Gemma 2 2B, Qwen 2 7B, and Llama 2 7B) to reduce refusal rates on morally sensitive prompts. The uncensored models show lower refusal rates across all six moral dimensions and better align with human preferences compared to their censored counterparts, though refusal rates do not drop to zero.
Evidence
interventional
Key metric
Refusal rates and preference decomposition scores for uncensored versus censored models (reported in radar plots in Appendix F)
Caveat
Jailbreaking does not eliminate all refusals, and the approach is limited to a specific uncensoring technique; findings may not generalize to other jailbreaking methods or other models.