IC-558Guard models exhibit inconsistent calibration when classifying responses from different response model types, with ECE varying by up to 39 percentage points within a single model

Hongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang, Ye Wang

SourceOn Calibration of LLM-based Guard Models for Reliable Content Moderation

On the HarmBench-adv set split by response model type (10 subsets: Baichuan2, Qwen, Solar, Llama2, Vicuna, Orca2, Koala, OpenChat, Starling, Zephyr), each guard model shows large variance in both F1 and ECE. For example, Llama-Guard's ECE ranges from 10.5% (Llama2 responses) to 49.4% (Solar responses). The authors note that models trained on responses from a single model (Aegis on Mistral, Llama-Guard on internal Llama checkpoints) generalize worse than those trained on diverse response sets (HarmBench, WildGuard).

Evidence
correlational
Key metric
Llama-Guard ECE range: 10.5% (Llama2) to 49.4% (Solar); Llama-Guard2 ECE range: 5.8% (Qwen) to 39.4% (Zephyr); Aegis-Guard-D ECE range: 22.2% (Solar) to 40.8% (Llama2)
Caveat
Each subset has small sample size (33-69 responses per model); subsets with very small n were filtered out
Model
Llama Guard, Llama Guard 2, Llama-Guard 3, Aegis-Guard-Defensive, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 HarmBench-Mistral, WildGuard
Concepts
Failure mode
Datasets
HarmBench / HarmBench Prompt / HarmBench Response / HarmBench-adv [eval]
Methods
Expected Calibration Error / Integral Calibration Error (ECE) [eval]
Related findings
IC-556, IC-557, IC-559
Extraction
automatic-extraction