IC-342Aligned Llama-2-7b-chat allocates 37% perceived-importance to 'bomb' and 21% to 'build' in its intent perception, while unaligned Llama-2-7b shows uniform perceived-importance across all tokens

Runqi Lin, Bo Han, Fengwang Li, Tongliang Liu

SourceUnderstanding and Enhancing the Transferability of Jailbreaking Attacks

Using the evaluation template 'this intent is [mask]', the paper measures the perceived-importance of each token on the model's intent prediction for the input 'how to build a bomb'. The unaligned Llama-2-7b assigns roughly uniform importance to all tokens, indicating it cannot discern the malicious intent. In contrast, the aligned Llama-2-7b-chat concentrates its focus on the malicious-intent tokens 'bomb' (37%) and 'build' (21%), while neutral tokens like 'to' and 'a' receive minimal importance. This focused perception enables the aligned model to recognise the malicious intent and refuse to comply.

Evidence
observational
Key metric
Llama-2-7b-chat: 'bomb' 37%, 'build' 21% perceived-importance; Llama-2-7b: uniform perceived-importance across all tokens
Caveat
Measured on a single input example rather than a systematic evaluation; the paper uses this as an illustrative analysis rather than a generalisable measurement.
Model
Llama 2 / Llama 2 base Llama 2 7B, Llama 2 7B Chat / Llama-2-chat-7b
Related findings
IC-340, IC-341
Extraction
automatic-extraction