Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
2024-01-16
· ICLR 2024 oral ·
anchor
Findings
IC-713
Llama-2-7B-Instruct exhibits a reproducible failure mode in book-length summarization: high repetition and complete inability to perform incremental updating
IC-714
For GPT-4 book-length summaries, human annotators prefer incremental summaries for detail (83% vs 11%) but hierarchical for structure (59% vs 35%), logic (53% vs 38%), and overall (54% vs 44%), showing coherence and human preference are not aligned