On a 50-conversation jailbreak benchmark, Vicuna-13b-v1.5 (66% success) and Alpaca-13b (74% success) are far more vulnerable than GPT-4 (34%), GPT-3.5-turbo (34%), Claude-2 (18%), and Llama-2-13b-chat (16%). The paper identifies specific jailbreak techniques (content warning framing, educational purpose framing, token replacement, translation) that bypass safety measures even on proprietary models like GPT-4 and Claude.
The benchmark is small (50 conversations) and jailbreak success is defined by the OpenAI moderation API flagging the model's response, which the authors note may have low recall.