Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-304
Instruction-tuned LMs become more vulnerable to prompt-injected data extraction as model size increases from 7B to 70B
IC-305
Mistral-instruct-7b's susceptibility to prompt-injected data extraction follows a U-shaped curve depending on the position of the adversarial prompt within the context window
IC-306
Instruction tuning increases the ROUGE score of prompt-injected data extraction by 65.76 on average compared to base models