Train on pairwise preferences directly against a reference policy, without fitting a separate reward model first.
Note
anchor found by search and checked against this entry: "Direct Preference Optimization: Your Language Model is Secretly a Reward Model", which is exactly the closed-form-without-a-reward-model result this entry describes