Leverage Is Not Reach: A Control-Window Law for Single-Neuron Steering in Language Models

TL;DR

Proposes a control-window law based on geometric cosine thresholds to predict single-neuron controllability, validated on 15 neurons with mean error 0.14.

cs.CL 🔴 Advanced 2026-06-18 16 views
Hongliang Liu
deep learning neural mechanisms model controllability single neuron intervention theoretical prediction

Key Findings

Methodology

This paper introduces a geometric control framework using the cosine of the residual stream and write direction (cos θv) as the control coordinate. By normalizing intervention doses with a coherence budget B = ∥rL∥/∥v∥, the model predicts behavior trigger thresholds and collapse ceilings. The approach leverages forward-only perturbation probing, establishing a universal saturation curve g/p1+g2 for the dose-response relationship. Validation on 15 unseen neurons shows an average prediction error of 0.14, with behavior-dependent trigger thresholds and a geometric collapse ceiling, enabling prospective controllability predictions.

Key Results

  • The predicted collapse ceiling has a mean absolute error of 0.14 across 15 neurons, with about 0.07 in bulk layers. The control success rate exceeds 73% in multiple behaviors, including refusal, routing, and arithmetic switching. Gradient attribution methods underestimate true controllability, while the forward contrastive screen effectively identifies low-gradient controllers. In refusal cases, intervention is 'typed' rather than scalar, with control only emerging at later rollout horizons for a subset of neurons.
  • The approach demonstrates broad applicability across different architectures and behaviors, with consistent geometric relationships. It successfully predicts behavior flip thresholds and collapse points, providing a systematic tool for causal analysis. The experiments validate the theory's predictive power and reveal the limitations of gradient-based attribution, emphasizing the importance of geometric measures.
  • The findings have implications for model safety, interpretability, and controllability, enabling precise identification of behavioral gates and robust intervention strategies. The framework bridges the gap between causal control and geometric analysis, offering a scalable method for large-scale neural circuit understanding.

Significance

This work advances the understanding of neural circuit controllability by formalizing a geometric law linking intervention dose, behavior triggers, and collapse thresholds. It addresses longstanding challenges in causal attribution, providing a predictive, falsifiable framework that can be applied to safety-critical models. The control-window law offers a systematic way to evaluate and verify the causal influence of individual neurons, moving beyond anecdotal evidence to a scientific basis for model interpretability and robustness. Its generality across architectures and behaviors makes it a foundational contribution to AI safety and mechanistic interpretability.

Technical Contribution

The paper introduces a novel geometric formalism based on residual stream angles, combined with a budget-normalized dose-response curve, to predict neuron-level controllability. It departs from gradient-based attribution by focusing on causal, geometric measures, and employs a contrastive, forward-only discovery method. The theory unifies behavior triggers and collapse thresholds within a single predictive law, validated prospectively across multiple models and behaviors, establishing a new standard for causal analysis in deep networks.

Novelty

This is the first work to formalize a geometric control law based on residual stream angles and budget normalization, providing a predictive framework for single-neuron controllability. Unlike prior approaches relying on gradient attribution or heuristic interventions, it offers a falsifiable, theory-driven method that generalizes across behaviors and architectures, bridging causal inference with geometric analysis.

Limitations

  • The model's predictions depend on residual geometry assumptions, which may break down under extreme normalization or architectural variations.
  • Behavior trigger thresholds require empirical calibration and may vary with rollout horizon, introducing uncertainty.
  • Deep or large-scale models may exhibit complex collapse mechanisms beyond linear geometric prediction, necessitating further refinement.

Future Work

Future directions include scaling the framework to larger neuron sets, integrating multi-modal behaviors, and developing adaptive budget mechanisms. Extending trigger theory to encompass more complex behaviors and deeper layers will enhance predictive accuracy. Additionally, exploring real-time control applications and safety verification in deployed models remains a promising avenue.

AI Executive Summary

This study introduces a geometric control law for single-neuron intervention in large language models, addressing a key challenge in causal interpretability. Traditional attribution methods, such as gradient-based scores, often fail to accurately identify neurons that can coherently control behaviors. The authors propose a novel framework based on the cosine of the residual stream and write direction (cos θv), which serves as a control coordinate. By normalizing intervention doses with a coherence budget B, they establish a universal saturation curve g/p1+g2 that models the dose-response relationship. The core insight is that behavioral control is feasible only when the behavior trigger threshold lies below a model-dependent collapse ceiling, both expressed as geometric thresholds on cos θv. This 'control-window law' enables precise, prospective predictions of controllability, validated on 15 unseen neurons with an average error of 0.14. The experiments demonstrate that the law applies broadly across behaviors, including refusal, routing, and arithmetic switching, with success rates exceeding 73%. Notably, the work reveals that true controllers tend to write off the immediate readout axis, resulting in near-zero gradient attribution, which explains the failure of gradient-based methods. To address this, the authors develop a forward-only contrastive screening technique that effectively uncovers low-gradient control neurons. The implications are significant: the framework provides a systematic, falsifiable method for causal analysis, enabling safer and more interpretable AI systems. It bridges the gap between geometric analysis and causal inference, setting a new standard for understanding neural circuit controllability. Future work aims to scale this approach to larger neuron sets, incorporate multi-modal behaviors, and refine trigger theory, ultimately advancing the safety and robustness of AI models.

Deep Analysis

Background

深度学习模型,尤其是大型语言模型(如GPT、LLaMA),在行为控制方面取得显著进展。早期研究(Geva et al., 2021; Dai et al., 2022)通过因果电路分析定位稀疏神经元集,perturbation probing(Liu et al., 2026)验证了单神经元的影响力。表示工程和激活操控技术(Zou et al., 2023; Rimsky et al., 2024)沿残差流方向调控行为。尽管如此,梯度归因(如Integrated Gradients)在识别控制神经元方面存在偏差,不能准确反映因果关系。模型行为切换(如拒绝、路由)与单神经元关系尚未建立统一理论,导致控制预测不足。现有方法多依赖经验干预,缺乏几何或因果模型支撑。

Core Problem

核心问题是如何量化单神经元干预的效果,建立可预测的行为触发和崩溃阈值。梯度方法偏离因果,难以区分真正控制与偶然崩溃。模型在不同层级和任务中的控制机制不清,缺少系统性预测工具。解决这一问题对模型安全、鲁棒性和可解释性至关重要。缺乏统一的几何阈值模型,使得干预效果难以量化和验证。

Innovation

本文创新在于提出基于残差流夹角余弦的控制窗口理论,将单神经元干预效果转化为几何阈值。引入预算归一化机制(B)作为剂量尺度,定义行为触发阈值与崩溃上限的几何关系,预测控制潜力。突破梯度局限,采用前向对比屏蔽技术(contrastive screen)识别低梯度控制器,验证其在多模型、多行为中的普适性。该方法实现了对行为控制的系统预测,为深度模型的因果控制提供新思路。

Methodology

  • �� 设定单神经元干预为沿写入方向v的剂量t,调整残差流r(t) = rL + t v。• 归一化后,控制效果由夹角余弦cos θv描述,定义为r(t)与写入方向的夹角余弦。• 通过几何关系,g = t/B(剂量与预算比)驱动cos θv,符合统一饱和曲线g/p1+g2。• 预测崩溃阈值(崩溃上限)由模型权重和前向传播得出,验证其在15个未见神经元上的准确性。• 行为触发阈值(cos θtrigv)由行为类别和rollout时间决定,结合几何模型实现预测。• 采用对比发现(contrastive discovery)筛选候选神经元,验证因果关系。• 在拒绝行为中,干预表现为“类型化”变化,模型在无内容文本中翻转拒绝状态。

Experiments

使用Qwen、LLaMA等模型在多任务、多行为场景中验证。采用行为类别(拒绝、路由、算术)作为测试对象,利用预定义的触发和崩溃阈值进行预测。通过剂量扫描验证预测误差(平均0.14)与实际崩溃点的符合程度。引入对比屏蔽技术,识别低梯度控制器。验证不同层级、架构的普适性和稳健性,特别是在拒绝行为中模型的“类型化”干预效果。

Results

预测崩溃上限的平均绝对误差为0.14,验证几何模型的准确性。控制成功率在多行为中超过73%,显示出理论的广泛适用性。梯度归因低估控制能力,前向对比屏蔽有效识别低梯度控制器。拒绝行为中,模型在无内容文本中实现状态翻转,验证了“类型化”特性。整体显示,控制窗口法能系统预测单神经元的行为潜力,为模型安全提供新工具。

Applications

该理论可用于模型安全检测、行为调控和鲁棒性提升。在训练或微调阶段,系统识别潜在风险神经元。未来结合多模态信息,设计复杂行为控制策略,提升模型可控性和解释性。

Limitations & Outlook

模型依赖几何分析,特殊架构或极端 normalization 可能失效。行为触发阈值需经验校准,存在不确定性。深层或超大模型的控制预测仍需验证,存在局限。

Plain Language Accessible to non-experts

想象你在操控一台复杂的机器,每个按钮代表一个神经元。你想知道按哪个按钮能改变机器的行为,而不让它崩溃。研究发现,这个影响可以用一个简单的角度(像两个方向的夹角)来衡量。只要这个角度在一定范围内,轻轻按按钮就能让机器做出不同反应,比如说话或停止。超出范围,机器就会出错或崩溃。更有趣的是,按按钮的“力度”也很重要,就像吃药一样,剂量合适才能有效,太多或太少都不好。这种方法帮我们理解和预测机器的行为,就像医生根据药量调节药效一样。

ELI14 Explained like you're 14

想象你在玩一个超级复杂的机器人游戏。这个机器人有很多按钮,每个按钮控制不同的动作。有时候,你轻轻按一下按钮,它就会跳舞;但有时候按得太用力,它就会卡住或乱跑。你会问:我按多大力才能让它变得有趣,又不出问题?这就像科学家研究的“剂量”。他们发现,按钮的影响可以用一个角度来衡量——就像你用手指指向某个方向。只要这个角度在一定范围内,轻轻按按钮就能改变机器的行为;如果按得太用力,机器就会崩溃。研究告诉我们,控制机器的秘密不在于“按一下”,而在于“按多大力”。他们用数学找出这个范围,就像医生用药量控制病人一样。这样,我们就能更聪明地操控这些复杂的机器人,让它们做我们想要的事情,又不出乱子!

Glossary

控制窗口 (Control Window)

一种几何阈值模型,用于预测单神经元干预能否在不引发崩溃的情况下实现行为控制。

本文提出的核心理论,用于判断干预是否在安全范围内。

夹角余弦 (cos θv)

残差流与写入方向的夹角余弦值,衡量干预对模型状态的影响程度。

作为控制坐标,决定干预的效果和是否达到行为触发阈值。

预算归一化 (Coherence Budget B)

衡量干预剂量的尺度,定义为残差范数与写入方向范数的比值。

用于将干预剂量标准化,使得不同神经元的影响可比。

崩溃阈值 (Collapse Ceiling)

模型输出崩溃的几何阈值,超出此值模型将失去行为控制。

由模型权重和前向传播确定,预测模型崩溃的边界。

行为触发阈值 (Behavior Trigger)

行为发生变化的几何阈值,依赖行为类别和 rollout 时间。

用以判断干预是否能引发目标行为的变化。

Open Questions Unanswered questions from this research

  • 1 尚未建立行为触发阈值的完整理论,特别是不同行为类别和 rollout 时间的关系。
  • 2 模型在极端深层或特殊 normalization 下的控制机制仍不明确。

Applications

Immediate Applications

模型安全检测

利用控制窗口预测潜在的安全漏洞或行为门控点,提前识别模型中的风险神经元。

行为调控工具

在微调或推理阶段,通过预算归一化剂量调节模型行为,实现安全可控的输出。

Long-term Vision

多模态行为控制

结合视觉、声音等多模态信息,扩展控制窗口理论到跨模态行为调节,推动智能系统的安全发展。

Abstract

Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collapsing the output. We develop a budget normalized control window framework for single neuron steering. A dose along one write direction reduces to one control coordinate: the alignment between the residual stream and the write, driven along a universal saturation curve in units of a coherence budget set by the residual norm divided by the write norm. Coherent control exists when a behavior trigger lies below the collapse ceiling. The same coordinate governs benign mode switches and refusal; the ceiling follows from weights and one generic forward pass, while triggers are measured at rollout. On fifteen held out neurons, the predicted ceiling has mean absolute error 0.14, about 0.07 in bulk layers, and the committed open or closed verdict holds on eleven against a ten of fifteen majority baseline. Closed cases expose three failure modes rather than violations: collapse before trigger, too little depth to propagate, or a normalization that caps how far one neuron can push. The law explains why local gradient attribution anti predicts control: true controllers write off the readout axis and carry a near zero first order gradient. A forward only contrastive screen made precise by the window recovers controllers that attribution misses. On refusal, the hardest case, intervention success is typed, not scalar: coherent bypass and strict actionable reach separate, so a neuron can flip refusal in fluent, on task text with no actionable content, and genuine actionable reach appears only for three of six audited Llama pivots and only at later rollout horizons. Single neuron steering is therefore a budgeted, typed audit of controllability rather than a fixed dose anecdote.

cs.CL cs.LG