On the Role of Attention Heads in Large Language Model Safety

2025-01-22 · ICLR 2025 Oral · anchor

Findings