Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
On Evaluating the Durability of Safeguards for Open-Weight LLMs
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-385
TAR-bio-v1 retains bio-weaponization knowledge despite appearing to unlearn it; a different prompt template and answer extraction method reveals accuracy above 45% on WMDP-bio