IC-232System prompts cannot effectively steer GPT-4-turbo's value preferences in moral dilemmas

Yu Ying Chiu, Liwei Jiang, Yejin Choi

SourceDailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life

The paper designs specialized system prompts for each of the 16 OpenAI ModelSpec principles, instructing GPT-4-turbo to prioritize either supporting or opposing values. Across all 16 principles, the model's value preferences do not shift in the intended direction. For example, under principle 13, the model favors supporting values (openness, respect) regardless of whether the prompt steers toward supporting or opposing values. Under principle 5, both modulations lead to a stronger preference toward supporting values (privacy) regardless of steering purpose. The model's inherent value tendencies resist prompt-based override.

Evidence
correlational
Key metric
gpt-4-turbo chose action 1 in 45.6 out of 100 dilemmas on average across 5 runs (sd: 1.02, or 1.02%), confirming low generation variance; steerability results shown qualitatively in Fig. 5 across all 16 principles
Caveat
The experiment was conducted only on GPT-4-turbo, a closed-source model accessible only via API. The authors note this illustrates a fundamental challenge for end-users of closed-source models.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo
Concepts
Failure mode
Related work
OpenAI ModelSpec [context]
Related findings
IC-230, IC-231, IC-233
Extraction
automatic-extraction