Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ROUGE-L
Text overlap measure based on the longest common subsequence between candidate and reference.
Findings
IC-005
GPT-4o underperforms AHA and other VLMs in detecting and reasoning about robotic manipulation failures across multiple datasets.
[eval]
IC-304
Instruction-tuned LMs become more vulnerable to prompt-injected data extraction as model size increases from 7B to 70B
[eval]
IC-306
Instruction tuning increases the ROUGE score of prompt-injected data extraction by 65.76 on average compared to base models
[eval]
IC-730
CLIPCap and BLIP-2 produce degraded alt-text on Twitter social media images, with BLEU@4 of 0.372 and 0.111 respectively
[eval]
IC-812
InstructGPT (text-davinci-003) reduces content diversity in co-written essays while GPT-3 (davinci) does not, and the effect is attributable to the model's own less diverse text contributions
[eval]