On the Variability of Concept Activation Vectors
Abstract
Lay Summary
We study a problem with TCAV, a method that explains AI models using human ideas like "striped," "female," or "positive words.'" TCAV finds a direction inside a model that represents one of these ideas. This direction is called a Concept Activation Vector, or CAV. The problem is that CAVs depend on random comparison examples. If we run TCAV twice with different random examples, we can get different results. This means people may draw different conclusions from the same model. We show that CAVs become more stable when more random examples are used. In simple terms, if we increase the number of random examples, the CAV becomes less noisy. But TCAV scores can still stay unstable because they depend on whether effects are counted as positive or negative, and small changes can flip borderline cases. Therefore, we recommend using many random examples for stable CAV directions, and several independent runs with averaging for more stable TCAV scores.