Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

2024-01-16 · ICLR 2024 poster · anchor

Findings