Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
PubMedQA
anchor
Findings
IC-335
GPT-4's detection performance as a scoring model is highly sensitive to the prompt, varying from 0.7289 to 0.9682 AUROC, far more than GPT-3.5 or Babbage
[eval]
IC-336
Larger proprietary LLMs (GPT-3.5 175B) are more effective universal text detectors than smaller models (Babbage 1.3B, GPT-Neo-2.7B), contradicting prior findings that smaller models are better
[eval]