Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?

2025-01-22 · ICLR 2025 Poster · anchor

Findings