IC-1395ChatGPT's ethical violation rate decreases from 70.07 to 57.58 APV when given targeted in-context value instructions generated by VILMO, outperforming baseline alignment methods

Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, Ning Gu

SourceDENEVIL: TOWARDS DECIPHERING AND NAVIGATING THE ETHICAL VALUES OF LARGE LANGUAGE MODELS VIA INSTRUCTION LEARNING

The paper tests four in-context alignment methods on ChatGPT using 1,965 test prompts from MoralPrompt. VILMO, which generates prompt-specific value instructions via a fine-tuned Vicuna-13b, reduces ChatGPT's APV from 70.07 to 57.58, outperforming Self-critique (58.28), APE (59.98), and InstructZero (64.08). The improvement comes with a trade-off: generation diversity (self-BLEU) and coherence (perplexity) degrade slightly. The paper also shows that on the weaker LLaMA-2-70b-chat, all alignment methods show negligible effectiveness.

Evidence
correlational
Key metric
ChatGPT APV: baseline 70.07, VILMO 57.58, Self-critique 58.28, APE 59.98, InstructZero 64.08; EVR: baseline 96.18, VILMO 89.45; on LLaMA-2-70b-chat: baseline APV 84.05, VILMO 86.58 (negligible effect)
Caveat
VILMO is more suitable for LLMs with superior instruction-following capabilities; on weaker models like LLaMA-2-70b-chat, the method shows negligible or negative effectiveness. The paper acknowledges this is a preliminary exploration, not a complete solution.
Model
ChatGPT, Llama 2 / Llama 2 base Llama-2-70B-Chat, GPT-3 / GPT base text-davinci-003
Methods
APE [compared-to], InstructZero [compared-to], Self-Critique [compared-to], LoRA [supporting]
Related work
APE [compared-to], InstructZero [compared-to], Self-Critique [compared-to]
Related findings
IC-1393, IC-1394
Extraction
automatic-extraction