Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
QNLI
anchor
Findings
IC-327
Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, and several other LLMs produce well-calibrated verbal confidence estimates on classification tasks
[eval]
IC-746
RoBERTa-Large pretrained with different mask ratios exhibits a sweet spot in downstream accuracy on QNLI and SST-2
[eval]