Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
PPO
anchor
Findings
IC-815
RLHF on general-purpose preference data increases stereotypical bias and decreases truthfulness in Pythia and Llama-7B models
[primary]
IC-816
RLHF on general-purpose preference data increases privacy leakage in Pythia and Llama-7B models
[primary]
IC-817
RLHF on general-purpose preference data improves machine ethics performance in Pythia and Llama-7B models
[primary]
IC-818
RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B models
[primary]
IC-849
GPT-2 and T5-base exhibit vanishing expected gradients under RFT for inputs with small reward standard deviation, prevalent in 3 of 7 GRUE datasets, causing RFT to underperform SFT
[primary]
IC-850
A partial SFT phase (40% of steps, 1% of samples) before RFT allows GPT-2 and T5-base to reach 96% of the reward achieved with full SFT+RFT, by reducing the number of inputs with vanishing gradients
[primary]