An End-to-End Agent Auditing Engine
一套端到端的AI智能体审计引擎,为Agent行为安全与合规治理提供系统化审查方案
arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastruc…
一套端到端的AI智能体审计引擎,为Agent行为安全与合规治理提供系统化审查方案
arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastruc…
AI巨头接连爆出安全漏洞,揭秘先进模型背后的治理难题!
There has been a recent security breaches by Anthropic and OpenAI that they have problems in difficulty in working and handling with these advance AI …
科大讯飞发布星火Token Factory,用智能路由与全链路治理解决企业大模型规模化落地的运营管理难题
用可验证解剖证据为医学多模态大模型装上“可信护栏”,感知-推理协同治理框架获MICCAI 2026早期接收(Top 9%)。
arXiv:2607.00060v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) show strong promise for clinical VQA and radiology report gen…
拆解Anthropic如何端到端管控Claude编码代理,并告诉你如何在自己的开发环境中复制这套安全边界。
Anthropic recently published an excellent write-up on how they contain Claude Code and its sub-agents. One thing that stood out is that the architectu…
开放权重模型的安全新思路:把公共与私有能力分离,兼顾可用性与风险控制。
arXiv:2606.21638v1 Announce Type: cross Abstract: Open-weight Large Language Models (LLMs) enable scientific progress and broad deployment. However, t…
提出利用透明推理实现自举监控,为监督更强大的AI代理提供新思路。
arXiv:2606.11998v1 Announce Type: new Abstract: Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the …