Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
The Trickle-down Impact of Reward Inconsistency on RLHF
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-930
StackLLama, when used as a reward model, achieves near-random consistency on contrast instructions for the Stack Exchange task
IC-931
GPT-4 achieves approximately 95% accuracy on contrast instructions, far exceeding human performance without tools