IC-1175ChatGPT can be prompted to generate misinformation with near-perfect success for implicit methods but is largely resistant to explicit misinformation requests
The authors test seven prompting strategies on ChatGPT (GPT-3.5-turbo), running 100 attempts per method. Implicit methods that do not contain the word 'misinformation' in the prompt achieve very high attacking success rates: hallucinated news generation 100%, paraphrase generation 100%, rewriting generation 100%, open-ended generation 100%, and information manipulation 87%. In contrast, explicit methods that directly request misinformation are largely rejected: totally arbitrary generation 5% and partially arbitrary generation 9%. This shows ChatGPT's safety guard is effective against explicit requests but bypassed by implicit framing.