1
Transformers converge to invariant algorithmic cores
一项新研究揭示Transformer在训练中会收敛到不变的算法核心,为理解模型行为与电路机制提供新视角。
arXiv:2602.22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the s…