TCAV (Testing with Concept Activation Vectors)

anchor

Trains a linear classifier separating activations of concept examples from random ones, then measures how far the model's output moves along the resulting normal vector.

Findings