1
A Mechanistic View of Authority Hierarchy in LLM Sycophancy
用机制视角拆解大模型谄媚行为中的权威层级影响,揭示模型如何对不同身份用户区别回应
arXiv:2607.00415v1 Announce Type: cross Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize …