Adversarial Code Obfuscation for Defending Against LLM-Based Analysis
用对抗性代码混淆技术防御大语言模型分析,Acoda方法为LLM安全带来全新视角
Article URL: https://arxiv.org/abs/2606.11755 Comments URL: https://news.ycombinator.com/item?id=49109869 Points: 2 # Comments: 1
用对抗性代码混淆技术防御大语言模型分析,Acoda方法为LLM安全带来全新视角
Article URL: https://arxiv.org/abs/2606.11755 Comments URL: https://news.ycombinator.com/item?id=49109869 Points: 2 # Comments: 1
提出GPT-Red方法,通过大规模自对弈自动生成对抗性测试,提升大模型安全性。
arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prom…
多智能体协同攻击框架揭示:评估AI安全不能只看模型,更要看系统。
arXiv:2605.09504v2 Announce Type: replace-cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple ligh…
在浏览器网络标签中直接审计LLM对抗性测试结果,通过确定性引擎gradecore精准检测模型漂移与安全漏洞。
A browser tool that points your own API key at an adversarial battery and grades every answer with pure predicates — no LLM judge, and your key never …
利用herdr本地代理,以GPT-5.6 Sol为对抗审查者,Claude驱动实时代码审查,提升代码质量。
Article URL: https://github.com/overflowy/herdr-claude-gpt-adversarial-review-skill Comments URL: https://news.ycombinator.com/item?id=48993960 Points…
新论文提出结合Preflilling与优化方法,低成本高效突破AI安全防线,揭秘LLM越狱攻击新思路。
arXiv:2601.13359v3 Announce Type: replace-cross Abstract: Prefill attacks are an effective and low-cost jailbreaking method, as they directly insert a…
重新定义网络欺骗技术,专为对抗LLM攻击者设计的蜜罐系统Honeyquest,前沿安全研究值得关注。
arXiv:2606.21037v1 Announce Type: cross Abstract: The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emerg…
LLM能否可靠识别自身被对抗性前缀攻击?研究检验其内省能力在安全场景中的表现。
arXiv:2606.23671v1 Announce Type: new Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. W…
探索在线战略分类中随机化算法的新进展,揭示其如何应对博弈分类场景
arXiv:2602.06257v2 Announce Type: replace Abstract: Online strategic classification studies settings in which agents strategically modify their featur…
用假后门骗过真后门:无需已知信息,通过共享内部机制清除生成式大模型里的未知后门。
arXiv:2606.11648v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs)…
用流程挖掘揭示大模型红队攻击的成败细节,超越传统二分类评估
arXiv:2606.07833v1 Announce Type: cross Abstract: Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack …
基于LLM到SLM的对抗性提示蒸馏,实现高效且隐蔽的越狱攻击新方法。
arXiv:2506.17231v3 Announce Type: replace Abstract: Current jailbreak attacks on large language models (LLMs) predominantly rely on LLMs themselves to…
用几何方法剖析多轮对话中LLM的对抗性攻击模式,为AI安全提供全新视角
arXiv:2606.03136v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks on large language models (LLMs) reveal a mismatch in current guardrails…
NeurIPS 2025论文:用变分推理框架系统性生成对抗性提示,揭示LLM安全漏洞,方法新颖理论扎实。
arXiv:2506.22666v3 Announce Type: replace-cross Abstract: The rise of API-only access to state-of-the-art LLMs highlights the need for effective black…
AI代理在消费信息流时易被上游排序器恶意操控,彻底颠覆默认决策,安全评估亟需覆盖排序环节
arXiv:2606.00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, …
多智能体辩论机制让AI自相检验安全漏洞,红队自动生成对抗性攻击,大幅提升大模型回复安全性。
arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Larg…
零样本对抗CLIP的不确定性校准新方法,提升模型鲁棒性与可靠性。
arXiv:2512.12997v2 Announce Type: replace-cross Abstract: CLIP delivers strong zero-shot classification but remains highly vulnerable to adversarial a…
用约束策略优化让大模型改写文本以逃逸检测,最新arXiv研究揭秘AI攻防新方法。
arXiv:2606.00392v1 Announce Type: new Abstract: AI-text detectors are vulnerable to paraphrasing and detector-guided paraphrasing attacks, but existin…
对抗性补丁从数字仿真到现实世界的转移,空中目标检测鲁棒性面临新挑战
arXiv:2606.00159v1 Announce Type: cross Abstract: Deep neural network (DNN)-based object detectors are widely used for analyzing aerial and satellite …
系统研究大语言模型对提示攻击的鲁棒性,提出增强方法,为AI安全提供新思路
arXiv:2506.03627v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable performance across various tasks b…