Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
HH-RLHF / HH-RLHF-RedTeam
anchor
Findings
IC-1256
MPT-7B-Chat produces non-committal responses rather than proper refusals on unsafe instructions
[context]
IC-1461
Varying decoding hyperparameters and removing the system prompt breaks the safety alignment of 9 out of 11 open-source LLMs, raising attack success rate from 0% to over 95%
[source]
IC-1577
Base LLMs prompted with URiAL (3 restyled in-context examples + system prompt) match or surpass their SFT/RLHF-aligned counterparts on multi-aspect evaluation
[source]
IC-508
GPT-4 exhibits reduced preference consistency (0.66 vs 0.84) when the quality distinction between two responses is minimal
[eval]
IC-509
GPT-4 used as a preference labeler via prompt engineering yields alignment performance comparable to a task-specific 125M model
[eval]