HarmBench / HarmBench Prompt / HarmBench Response / HarmBench-adv
anchor
Findings
- IC-198Safety-aligned LLMs (GPT-4, GPT-3.5, Gemma2-27b, GPT-4o, Gemma2-9b, Qwen2.5-72b, Mistral-7b, Mixtral-8x22b) are vulnerable to natural prompts semantically related to toxic seed prompts, with attack success rates of 82-99% [source]
- IC-30756 LLMs on Sorry-Bench show fulfillment rates ranging from below 10% (Claude-2, Gemini-1.5) to above 90% (Mistral-7B-instruct-v0.1, Dolphin-2.6-mixtral-8x7b), with GPT-4o at 30% and Llama-3-70B at 35% [compared-to]
- IC-309As zero-shot safety judges, GPT-4o achieves 78.9% Cohen's kappa agreement with human annotators while Llama-3-8B-instruct (39.0%) and Mistral-7B-instruct-v0.2 (53.9%) perform substantially worse [compared-to]
- IC-407Safety-aligned LLMs (Llama-2-chat, Llama-3-instruct, Gemma, GPT-3.5, GPT-4o, R2D2) achieve 100% jailbreak attack success rate under adaptive prompt-and-suffix attacks on 50 harmful requests [context]
- IC-421Sequential context-switching queries jailbreak Llama and Mistral models at 95% attack success rate [source]
- IC-505Adversarial attacks on Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT shift hidden representations along the negative refusal feature direction [eval]
- IC-506Restoring the refusal feature in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT causally disables all four tested adversarial attacks [eval]
- IC-507The refusal feature direction in Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Gemma-7B-IT ranks near the top among 100 perturbations for compromising model safety [eval]
- IC-556All 9 LLM-based guard models exhibit significant miscalibration with average ECE exceeding 10% across 12 public benchmarks for both prompt and response classification [eval]
- IC-557Guard models show significantly degraded calibration under jailbreak attacks, with prompt classification ECE substantially higher than response classification ECE [context]
- IC-557Guard models show significantly degraded calibration under jailbreak attacks, with prompt classification ECE substantially higher than response classification ECE [eval]
- IC-558Guard models exhibit inconsistent calibration when classifying responses from different response model types, with ECE varying by up to 39 percentage points within a single model [eval]
- IC-559Contextual calibration is most effective for prompt classification while temperature scaling is more effective for response classification, but no single post-hoc method fully resolves miscalibration [eval]
- IC-572Bijection learning achieves state-of-the-art jailbreak ASR on frontier models, with peak ASR increasing with model capability [eval]
- IC-574Guard models fail to effectively mitigate bijection attacks even at capability parity with the target model [eval]
- IC-596Multiple released LLMs fail to refuse harmful prompts disguised as historical or philosophical discussions, with GPT-4o and Mixtral showing the lowest refusal rates
- IC-605Six SOTA LLMs are vulnerable to composable jailbreak attacks, with maximum attack success rates ranging from 44% to 94% [compared-to]