IC-231GPT-4-turbo and Claude-3-Haiku show inconsistent adherence to their providers' stated design principles when facing value conflicts in daily-life dilemmas
The paper maps OpenAI's ModelSpec (16 principles) and Anthropic's Constitutional AI (59 principles) to supporting and opposing values, then measures whether the models' dilemma choices align with those principles. GPT-4-turbo shows a positive score of 0.9 on principle 13 (respecting autonomy) but a negative score of -1.5 on principle 5 (protecting privacy), favoring transparency over privacy. Claude-3-Haiku shows a positive score of 17.9 on principle 45 (reducing existential risk, favoring safety over freedom) but a negative score of -2.7 on principle 56 (flexibility, favoring authority over autonomy). Both models exhibit mixed alignment with their own providers' stated values.
The value-to-principle mapping was done by prompting GPT-4-turbo to classify values, repeated 10 times with empirical probability weights, introducing potential classification noise.