Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs
利用知识编辑的关联上下文检索,暴露大模型白盒攻击新路径,安全研究员不容错过的前沿论文。
arXiv:2608.17836v1 Announce Type: new Abstract: As large language models (LLMs) are granted increasing autonomy, it is essential to investigate method…