Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Lawma: The Power of Specialization for Legal Annotation
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-096
Large commercial and open-weight models achieve 70-78% accuracy on CASELAWQA, with Claude 3.7 Sonnet at the top
IC-097
Chain-of-thought prompting outperforms few-shot direct QA for Llama 3 models above 8B parameters, while few-shot is best below 3B
IC-098
LegalBERT performs below the constant classifier baseline on CASELAWQA due to its 512-token context window