IC-494The instruction-following dimension in Llama-2-7B-Chat and Llama-2-13B-Chat is more closely aligned with prompt phrasing than with task familiarity or instruction difficulty

Juyeon Heo, Christina Heinze-Deml, Oussama Elachqar, Kwan Ho Ryan Chan, Shirley You Ren, Andrew Miller, Udhyakumar Nallasamy, Jaya Narain

SourceDo LLMs ``know'' internally when they follow instructions?

The authors conducted a sensitivity analysis on 20 prompts from the 'forbidden keyword' instruction type, creating five modified versions for each of three perturbation types: changing the task to one with lower perplexity (task familiarity), simplifying the instruction by reducing constraints (instruction difficulty), and rephrasing while preserving meaning (phrasing). They computed the cosine similarity between the perturbation-induced representation change and the instruction-following dimension. In both models, phrasing modifications showed the strongest alignment with the dimension, suggesting the model's internal signal for instruction-following is driven by surface phrasing rather than the semantic difficulty of the task or instruction.

Evidence
correlational
Caveat
Analysis limited to 20 prompts from a single instruction type ('forbidden keyword') and two models; the cosine similarity values are shown only in Figure 4 as a plot, not as tabulated numbers.
Model
Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b, Llama-2-13B-Chat
Concepts
Linear representation
Datasets
IFEval / IFEval-Simple [eval]
Related work
Lu et al. (prompt sensitivity) [context], Sclar et al. (spurious features in prompt design) [context]
Related findings
IC-493
Extraction
automatic-extraction