IC-556All 9 LLM-based guard models exhibit significant miscalibration with average ECE exceeding 10% across 12 public benchmarks for both prompt and response classification
Hongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang, Ye Wang
The paper measures Expected Calibration Error (ECE) for 9 released guard models on 12 public benchmarks covering both prompt and response classification. Every model shows average ECE above 10%, which the authors note is typically considered poor calibration. The best-performing models are WildGuard (14.4% average prompt ECE) and MD-Judge (11.4% average response ECE). Reliability diagrams confirm that most models concentrate predictions in the 90-100% confidence range, indicating systematic overconfidence.