IC-335GPT-4's detection performance as a scoring model is highly sensitive to the prompt, varying from 0.7289 to 0.9682 AUROC, far more than GPT-3.5 or Babbage
The paper ablates five manually drafted prompts (from empty to complex) and measures the AUROC of fast-detectgpt when using different scoring models. GPT-4 shows the largest variation, with detection accuracy ranging from 0.7289 (empty prompt) to 0.9682 (prompt4), a spread of nearly 0.24. GPT-3.5 is less sensitive, ranging from 0.9071 to 0.9589, and Babbage is the least affected, ranging from 0.8921 to 0.9299. This indicates that GPT-4's predictive distributions shift substantially with the input context, making it the most prompt-dependent scoring model among those tested.