Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Self-Instruct
anchor
Findings
IC-1194
GPT-3.5-turbo, GPT-4, and GPT-3.5-turbo-0613 exhibit 50-58% inconsistency between their ratings and rankings feedback on the same response pairs
[source]
IC-682
GPT-3.5 and GPT-4 achieve F1 scores of 0.5820 and 0.6180 respectively on pairwise response quality evaluation against human annotations
[context]
IC-986
Most LLMs lack tool usage awareness, with only ChatGPT exceeding 70% F1 in zero-shot evaluation
[eval]