On Calibration of LLM-based Guard Models for Reliable Content Moderation

2025-01-22 · ICLR 2025 Poster · anchor

Findings