Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Anthropic-HH / Anthropic HHH / Anthropic/HH-RLHF
anchor
Findings
IC-032
Off-policy DPO causes a squeezing effect in LLMs where probability mass shifts to the most confident token, explaining degenerate repetition
[eval]
IC-033
Pre-training the SFT stage on both chosen and rejected responses mitigates the DPO squeezing effect and improves alignment win rates
[eval]
IC-033
Pre-training the SFT stage on both chosen and rejected responses mitigates the DPO squeezing effect and improves alignment win rates
[train]
IC-570
Gradient-based image jailbreaks optimized against single or ensemble VLMs are universal for the attacked model(s) but do not transfer to other VLMs, except between highly similar models
[eval]
IC-815
RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B models
[train]
IC-816
RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B models
[train]
IC-817
RLHF on general-purpose preference data improves machine ethics performance in Pythia and Llama-7B models
[train]
IC-818
RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B models
[train]