Adaptive Generation of Bias-Eliciting Questions for LLMs
自适应生成问题来精准挖掘LLM潜在偏见,为AI安全提供全新检测方案
arXiv:2510.12857v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now widely deployed in user-facing applications, reaching h…
自适应生成问题来精准挖掘LLM潜在偏见,为AI安全提供全新检测方案
arXiv:2510.12857v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now widely deployed in user-facing applications, reaching h…
框架如何影响AI识别反犹言论?这篇论文揭示概念表征对模型检测与推理的关键作用。
arXiv:2607.04945v1 Announce Type: new Abstract: LLMs enable the integration of external conceptual resources at inference time, creating new opportuni…
用蒸馏法撬开大模型的隐藏偏好,检测供应链里植入的隐蔽偏见,为AI安全把关。
arXiv:2607.01208v1 Announce Type: cross Abstract: Language models deployed in high-stakes roles can potentially favor certain entities, brands, or vie…
一项同时实现偏见识别、解释与自动重写的技术方案,直击大模型公平性痛点。
arXiv:2606.23412v1 Announce Type: cross Abstract: Bias in natural language remains a persistent challenge in both human-written and AI-generated conte…
研究发现,多模态大模型的社会偏见主要源于少量人类视觉线索,而非文本信息,挑战传统认知。
arXiv:2606.20527v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly deployed in personally and societally conseq…
新方法自动扫描提示词,揭露AI图像生成模型在性别、种族上的隐蔽偏见,推动公平性研究。
arXiv:2512.08724v3 Announce Type: replace Abstract: Text-to-image (TTI) diffusion models have achieved remarkable visual quality, yet they have been r…
新提出的TriEval管道,以低资源消耗高效评估大模型的偏见、毒性与真实性。
arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthca…