Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MM-SafetyBench
anchor
Findings
IC-021
Vision-language adaptation degrades safety in Llama-2-chat-7b even when training data is filtered for safety
[eval]
IC-022
Safety layers in Llama-2-chat-7b show substantial divergence during VL adaptation, correlating with safety degradation
[eval]
IC-571
Open-source VLMs (LLaVA, MiniGPT-4, InstructBLIP) are substantially more vulnerable to multimodal jailbreak attacks than Gemini-1.5-flash, with BAP attack ASR of 58–62% versus 40–41%
[eval]