Vanishing Gradients in Reinforcement Finetuning of Language Models

2024-01-16 · ICLR 2024 poster · anchor

Findings