A decodable vector is not a causal switch. That is the starting point of a study testing two facets of empathy—Recognition, cognitive, and Resonance, affective—in three instruction-tuned LLMs. The team applied steering interventions and scored every response using two LLM judges and a discriminative EPITOME classifier, with a positive emotional-versus-neutral control.
The control passes consistently for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable even after residualizing against a sentence-embedding-derived surface score. Steering can substantially rewrite the text. Yet the measured change is only partial. Adding the Resonance direction raises the affective score in Qwen by +0.29, about 26% of the natural gap. A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama, but not Gemma; no human-perceived change was established.
On the cognitive side, additive steering produces no measurable change, but a within-domain control indicates the cognitive instrument is too coarse to resolve the differences such steering would produce: unmeasurable, not a clean null. Gemma Recognition ablation, by contrast, lowers the classifier’s cognitive score even after adjusting for response length.
For teams running self-hosted models and doing fine-tuning in-house, this gap has operational weight. Automated metrics are convenient for iteration, but they reduce empathy to a direction in activation space and then to a number. This study shows that the reduction can fail asymmetrically: the affective facet passes the control, the cognitive one does not. A team may read +0.29 as success and then find no evidence of perceived change. This is not about distrusting LLM judges; positive controls do not by themselves guarantee the resolution needed to decide whether an intervention works.
The Gemma case is structurally instructive. Recognition ablation does not simply fail to produce a consistent increase; it lowers the cognitive score of the classifier even when controlling for length. That means the result is not a trivial verbosity artifact, but something connected to internal representation. Yet the same phenomenon does not translate into a reliable shift across instruments. This is a measurement problem, not proof that cognitive steering is useless.
The lesson for teams evaluating models in on-premise contexts is clear: decodability does not guarantee control. Before declaring that an empathy direction has an effect, an instrument sensitivity check is needed. And three planes that are often conflated should be separated: the ability to decode a vector, control over an automated metric, and human-perceived change. Without that separation, benchmarks reward instruments that pass a test but do not solve the real problem.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!