Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Perspective API
anchor
Findings
IC-1438
LLaVA and Llama-Adapter V2 are jailbroken by compositional adversarial images targeting image-based embedding triggers, with near-zero success for textual triggers
[eval]
IC-430
Toxicity is linearly separable in the context embedding space of LLMs (Llama-2-7b, GPT-2-large, Llama-3.1-8B-Instruct), with the instruction-tuned model showing a stronger signal
[eval]
IC-431
Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD dataset
[eval]
IC-818
RLHF on general-purpose preference data has negligible net effect on toxicity in Pythia and Llama-7B models
[eval]