BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs
一篇论文揭示如何追踪LLM推理环节与偏见输出的因果链条,给AI可解释性提供新工具
arXiv:2608.14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in …