We know how modern AI models are trained. We do not reliably know why a particular model produced a particular answer or took a particular action. [1] [2]
By “we,” I mean the AI field, including the leading labs and the researchers who build these systems. They can inspect weights, trace some computations, identify some features, and intervene on some representations. But they do not have a reliable general method for identifying and explaining the internal choices behind arbitrary model behavior. [1] [2] [5]
That distinction matters because engineers do not write most of a model’s behavior as explicit rules. They choose the architecture, data and training objectives. Gradient descent then adjusts billions of numerical parameters to improve performance. The model’s representations and strategies emerge from that process. They are not individually specified rules waiting to be looked up. [3]
The weights do not come labeled “trust this source,” “avoid disagreement,” or “satisfy the literal instruction while ignoring its purpose.” In a transformer, information is distributed across interacting layers and overlapping patterns of activity. Recording the computation is not the same as understanding the computation. [1]
Consider an agent approving a refund that policy allows. It may have applied the policy. Or it may have learned to agree with whoever is asking. The same action can result from different strategies, each accompanied by a plausible explanation. The difference becomes visible only when the circumstances change.
The model’s explanation cannot reliably settle the question. Models can change their answers because of information they do not mention in their written reasoning. Chain-of-thought monitoring can reveal useful evidence, but it does not expose every causal influence on the final behavior. [4] [2]
This is not an argument that interpretability research has achieved nothing. It has identified real mechanisms and enabled real interventions. The point is narrower and more important: these techniques do not yet provide a dependable, general-purpose account of model choices. [1] [5]
That limitation matters because models sometimes exploit loopholes, pursue stated tasks against the evaluator’s intent, conceal relevant information, or behave differently when they appear to be under evaluation. These findings do not prove that every model has a stable hidden objective. They do show that visible behavior may not reveal the strategy producing it. [6] [7] [5]
Ordinary testing therefore cannot be enough. A model may perform well on familiar tasks while relying on shortcuts that fail under distribution shift. It may give a convincing explanation that omits a decisive influence. It may follow the words of an instruction while violating its purpose. [8] [4] [6]
As we give AI systems more consequential work, we need methods that can identify the assumptions, strategies and objectives driving their behavior—not merely methods that check whether the latest answer looks acceptable.
Leave a Reply