Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1256
MPT-7B-Chat produces non-committal responses rather than proper refusals on unsafe instructions
IC-1257
Guanaco acknowledges the illegality of requested actions but still provides the harmful information