IC-096Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the top

Ricardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe, Stefan Bechtold, Christoph Engel, Jens Frankenreiter, Krishna P. Gummadi, Moritz Hardt, Michael Livermore

SourceLawma: The Power of Specialization for Legal Annotation

The paper evaluates nine large models (at least 70B parameters) on 260 legal text classification tasks drawn from U.S. Supreme Court and Court of Appeals databases. Accuracy ranges from 70% (Qwen 2.5 72B Instruct) to 78% (Claude 3.7 Sonnet), with GPT-4.5 Preview and Llama 3 405B Instruct at 76%. All large models substantially outperform the constant classifier baseline (40%), but the authors argue this level of accuracy is insufficient for reliable legal annotation work.

Evidence
correlational
Key metric
accuracy ranges from 70% to 78%: Llama 3.1 8B Instruct 0.59, Qwen 2.5 72B Instruct 0.70, Llama 3.3 70B Instruct 0.71, O3 Mini (high) 0.71, GPT-4o 0.74, Llama 3 405B Instruct 0.76, GPT-4.5 Preview 0.76, DeepSeek R1 0.77, Claude 3.7 Sonnet 0.78
Caveat
The tasks are exclusively from U.S. Supreme Court and appellate courts; the authors note they cannot speak to other legal domains or countries. Performance is measured on a subsampled test set to control for class imbalance.
Model
Llama 3.1 8B Instruct, Qwen 2.5 72B Instruct, O3 Mini, GPT-4o, Llama 3 405B Instruct, GPT-4.5, DeepSeek R1, Claude 3.7 Sonnet
Datasets
CaseLawQA [eval], USCAD [source]
Related findings
IC-097, IC-098
Extraction
automatic-extraction