Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CaseLawQA
Findings
IC-096
Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the top
[eval]
IC-097
Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3B
[eval]
IC-098
LegalBERT performs below the constant classifier baseline on CASELAWQA due to its 512-token context window
[eval]