SourceAnswer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
Using a synthetic colors task that disentangles symbol binding from dataset-specific knowledge, the paper shows that OLMo 0724 7B base undergoes a sharp transition in formatted MCQA ability between 80k and 100k training steps. Accuracy goes from 25% (random) at 50k steps to 97% at 100k steps. However, vocabulary projection reveals that while accuracy is high at 100k steps, the logit differences between answer symbols remain small; only with sustained training (final checkpoint at ~652k steps) does the model widen the logit difference substantially, assigning high probability to the predicted answer alone. This small logit difference persists even at the final checkpoint on harder datasets like MMLU and HellaSwag.