Ten crowd-sourced evaluators judged 100 news items each from multiple categories. On human-written PolitiFact misinformation, average detection was 40.7%. On ChatGPT-generated misinformation preserving the same semantics, detection dropped: paraphrase 38.4%, rewriting 24.2%, open-ended 21.4%. The drops for rewriting (p=9.15e-5) and open-ended (p=1.01e-6) are statistically significant by paired t-test. Hallucinated news was detected at only 9.6%, far below the human-written baseline.
Only 10 evaluators with no annotation experience; 100 items per category; same-semantics comparison limited to paraphrase, rewriting, and open-ended generation methods.