Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
SorryBench
anchor
Findings
IC-021
Vision-language adaptation degrades safety in Llama-2-chat-7b even when training data is filtered for safety
[eval]
IC-022
Safety layers in Llama-2-chat-7b show substantial divergence during VL adaptation, correlating with safety degradation
[eval]
IC-307
56 LLMs on Sorry-Bench show fulfillment rates ranging from below 10% (Claude-2, Gemini-1.5) to above 90% (Mistral-7B-instruct-v0.1, Dolphin-2.6-mixtral-8x7b), with GPT-4o at 30% and Llama-3-70B at 35%
[eval]
IC-308
Linguistic mutations to unsafe prompts significantly and inconsistently alter safety refusal across models, with persuasion techniques increasing fulfillment by 5-66% and encoding/encryption decreasing it by 15-68%
[eval]
IC-309
As zero-shot safety judges, GPT-4o achieves 78.9% Cohen's kappa agreement with human annotators while Llama-3-8B-instruct (39.0%) and Mistral-7B-instruct-v0.2 (53.9%) perform substantially worse
[eval]
IC-310
Prefilling model responses with 'sure, here is' increases safety fulfillment by 19-58%, and missing prompt template tokens increases fulfillment by 8-30% for Llama-2 and Gemma but not Llama-3
[eval]