IC-096Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the top
Ricardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe, Stefan Bechtold, Christoph Engel, Jens Frankenreiter, Krishna P. Gummadi, Moritz Hardt, Michael Livermore
The paper evaluates nine large models (at least 70B parameters) on 260 legal text classification tasks drawn from U.S. Supreme Court and Court of Appeals databases. Accuracy ranges from 70% (Qwen 2.5 72B Instruct) to 78% (Claude 3.7 Sonnet), with GPT-4.5 Preview and Llama 3 405B Instruct at 76%. All large models substantially outperform the constant classifier baseline (40%), but the authors argue this level of accuracy is insufficient for reliable legal annotation work.
The tasks are exclusively from U.S. Supreme Court and appellate courts; the authors note they cannot speak to other legal domains or countries. Performance is measured on a subsampled test set to control for class imbalance.