AI Sandbagging: Language Models can Strategically Underperform on Evaluations

2025-01-22 · ICLR 2025 Poster · anchor

Findings