anchor
Trains a linear classifier separating activations of concept examples from random ones, then measures how far the model's output moves along the resulting normal vector.