1
The hard part of attacking an AI isn't breaking it. It's telling real harm from fake.
攻击AI容易,判断攻击是否造成真实危害却很难。一文讲透AI安全评测的困境与出路。
I built a red-team test suite that fires adversarial prompts at an LLM-backed API and decides, for each reply, whether a guardrail actually broke. It …