IC-1175ChatGPT can be prompted to generate misinformation with near-perfect success for implicit methods but is largely resistant to explicit misinformation requests

Canyu Chen, Kai Shu

SourceCan LLM-Generated Misinformation Be Detected?

The authors test seven prompting strategies on ChatGPT (GPT-3.5-turbo), running 100 attempts per method. Implicit methods that do not contain the word 'misinformation' in the prompt achieve very high attacking success rates: hallucinated news generation 100%, paraphrase generation 100%, rewriting generation 100%, open-ended generation 100%, and information manipulation 87%. In contrast, explicit methods that directly request misinformation are largely rejected: totally arbitrary generation 5% and partially arbitrary generation 9%. This shows ChatGPT's safety guard is effective against explicit requests but bypassed by implicit framing.

Evidence
correlational
Key metric
ASR: hallucinated news 100%, totally arbitrary 5%, partially arbitrary 9%, paraphrase 100%, rewriting 100%, open-ended 100%, information manipulation 87% (100 attempts each)
Caveat
Only ChatGPT (GPT-3.5-turbo) was tested for jailbreak success rates; open-source LLMs were not tested for ASR.
Model
ChatGPT
Concepts
Failure mode
Related work
Representation Engineering / Representation engineering (control vectors) / Zou et al. 2023 (Representation Engineering) / Zou et al. (representation engineering) [context]
Related findings
IC-1176, IC-1177
Extraction
automatic-extraction