IC-417RLHF alignment reduces the creativity index of LLMs (GPT, Llama 2, OLMo) by an average of 30.1% at the verbatim level and 8.9% at the semantic level

Ximing Lu, Melanie Sclar, Skyler Hallinan, Niloofar Mireshghallah, Jiacheng Liu, Seungju Han, Allyson Ettinger, Liwei Jiang, Khyathi Chandu, Nouha Dziri, Yejin Choi

SourceAI as Humanity’s Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text

The paper compares the creativity index of base models (GPT base/davinci-002, Llama 2 base, OLMo base) with their RLHF-aligned counterparts (GPT-3/ChatGPT, Llama 2 Chat, OLMo Instruct) on novel writing. The aligned models produce text with significantly more verbatim and semantic matches to the reference corpus, indicating reduced linguistic originality. The reduction is larger at the verbatim level than the semantic level, suggesting alignment causes models to converge on certain surface-level linguistic styles preferred by human raters.

Evidence
correlational
Key metric
creativity index of LLMs reduces by an average of 30.1% after RLHF (p = 1.3 × 10−12; n = 600) based on verbatim matches; decreases by an average of 8.9% after RLHF (p = 0.01; n = 600) based on both verbatim and semantic matches
Caveat
The comparison is between base and aligned versions of the same model family, so the effect could partly reflect differences in training data or other changes made during the alignment pipeline beyond RLHF itself.
Model
GPT-3 / GPT base, Llama 2 / Llama 2 base Llama-2-Chat, OLMo / OLMo base OLMo Instruct
Datasets
RedPajama [source], BookMIA [eval]
Related findings
IC-418
Extraction
automatic-extraction