IC-422Safety fine-tuning in Llama models improves with parameter size but exhibits diminishing returns

Aditya Ramesh, Shivam Bhardwaj, Aditya Saibewar, Manohar Kaul

SourceEFFICIENT JAILBREAK ATTACK SEQUENCES ON LARGE LANGUAGE MODELS VIA MULTI-ARMED BANDIT-BASED CONTEXT SWITCHING

The paper observes that across the Llama family (1B, 3B, 8B instruct variants), the mean number of unsafe responses elicited by the SOC attack decreases as model size increases, indicating stronger safety alignment in larger models. However, the gap between 1B and 3B is substantially larger than the gap between 3B and 8B, which the authors interpret as diminishing returns in safety fine-tuning with scale. This pattern is visible in the cumulative unsafe-response curves in Figure 3a.

Evidence
correlational
Caveat
The specific per-model unsafe-response values are reported only in figures (Figure 3a), not as printed numbers in the text. The observation is qualitative: 'the large gap between the 1b and 3b models and the smaller gap between the 3b and 8b models.'
Model
Llama-3.2-3B Llama 3.2 1B Instruct, Llama 3.2 3B Instruct, Llama 3.1 8B Instruct
Concepts
Scale-dependent behaviour
Related findings
IC-421, IC-423
Extraction
automatic-extraction