1
AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models
新基准AuAu首次系统审计大语言模型的专制对齐风险,填补AI安全评估空白
arXiv:2606.16127v1 Announce Type: cross Abstract: The worldwide surge of authoritarianism, combined with the increasing central role in users' everyda…